Maik Fröbe

dblp:256/9118 · DBLP profile ↗
← Back
56ranked-venue papers in the field
17as first author
51since 2021 · last 2026
0000-0002-1003-981XORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 54 (17 first)Data Mining & Knowledge Discovery · 2
YearPublicationVenuePosition
2026 Overview of PAN 2026: Voight-Kampff Generative AI Detection, Text Watermarking, Multi-author Writing Style Analysis, Generative Plagiarism Detection, and Reasoning Trajectory Detection
Janek Bevendorff, Maik Fröbe, André Greiner-Petter, Andreas Jakoby, Maximilian Mayerl, Preslav Nakov, Henry Plutz, Martin Potthast, Benno Stein 0001, Minh Ngoc Ta, Yuxia Wang 0003, Eva Zangerle
ECIR (4)2
2026 Evaluating Information Retrieval Models Along Time: The LongEval Lab at CLEF 2026
Timo Breuer 0002, Matteo Cancellieri, Alaa El-Ebshihy, Maik Fröbe, Petra Galuscáková, Lorraine Goeuriot, Gabriel Iturra-Bocaz, Jüri Keller, Petr Knoth, Andreas Konstantin Kruff, Philippe Mulhem, Florina Piroi, David Pride, Philipp Schaer, Didier Schwab
ECIR (4)4
2026 The Third International Workshop on Open Web Search (WOWS)
Laura Caspari, Maik Fröbe, Sebastian Heineking, Michael Granitzer, Gijs Hendriksen, Djoerd Hiemstra, Martin Potthast, Arjen P. de Vries, Saber Zerhoudi
ECIR (3)2
2026 Evaluating the Efficiency and Effectiveness of Learned Sparse Retrieval with the lsr_benchmark
Maik Fröbe, Ferdinand Schlatt, Cosimo Rulli, Tim Hagen, Jan Heinrich Merker, Gijs Hendriksen, Carlos Eduardo Rosar Kós Lassance, Franco Maria Nardini, Rossano Venturini, Martin Potthast
ECIR (4)1
2026 Overview of Touché 2026: Argumentation Systems - Extended Abstract
Johannes Kiesel, Marc Feger, Tim Hagen, Sebastian Heineking, Maximilian Heinrich, Maik Fröbe, Katarina Boland, Wilhelm Pertsch, Julia Romberg, Ines Zelch, Stefan Dietze, Matthias Hagen, Martin Potthast, Benno Stein 0001
ECIR (4)6
2026 Talmud-IR: A Talmud-Inspired Interface for Discussing RAG Response Quality
Wojciech Kusa, Niklas Deckers, Maik Fröbe, Laura Dietz, Birte Platow, Mark Sanderson
ECIR (4)3
2026 Auto-Judge: A Cross-Task Benchmark for Comparing LLM Judges for Citation-Grounded RAG Systems
abstract
We present the Auto-Judge resource for the meta-evaluation of automated LLM judges, especially judges that evaluate Retrieval-Augmented Generation (RAG) systems that ground their response with citations. The resource couples (i) a data release of topics, pooled RAG responses, and human judgments, with (ii) a standardized protocol and software infrastructure for implementing "LLM-as-a-judge" methods in a reproducible and extensible way, including support for parameter sweeps and variant tracking.
Naghmeh Farzi, Tim Hagen, Eugene Yang 0001, Maik Fröbe, Ronak Pradeep, Hossein A. Rahmani, Xi Wang 0012, Oleg Zendel, Martin Potthast, Laura Dietz
SIGIR4
2026 ReNeuIR at SIGIR 2026: The Fifth Workshop on Reaching Efficiency in Neural Information Retrieval
abstract
The lack of efficiency in neural information retrieval remains one of the primary obstacles to deploying neural retrieval models as a first-stage retriever at scale. While recent tools have improved the standardized measurement of model efficiency, substantial progress is still needed to enable systematic comparative evaluation, for example, in terms of standards for systems and hardware configurations, cloud-based evaluation, benchmarks, and reproducibility. Beyond measurement, the IR community also needs stronger incentives to move in this direction, such as cost-efficiency as a review criterion or as efficiency and/or effectiveness measures in shared tasks, related teaching materials, efficiency-oriented user studies, and specialized awards for efficiency achievements. In particular, developing more efficient variants of highly effective retrieval algorithms should become an admissible research goal for PhD students if cost-efficiency is to become a first-class design objective in~IR. With ReNeuIR, we have established a recurring forum where these questions and new ideas are discussed and where the community comes together to collaboratively evaluate and improve efficiency benchmarking frameworks---most notably through the organization of a shared task focused on efficiency and reproducibility.
Maik Fröbe, Tim Hagen, Franco Maria Nardini, Martin Potthast
SIGIR1
2026 Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation
abstract
Tip-of-the-Tongue (ToT) retrieval benchmarks have largely focused on English, limiting their applicability to multilingual information access. In this work, we construct multilingual ToT test collections for Chinese, Japanese, Korean, and English, using an LLM-based query simulation framework. We systematically study how prompt language and source document language affect the fidelity of simulated ToT queries, validating synthetic queries through system rank correlation against real user queries. Our results show that effective ToT simulation requires language-aware design choices: non-English language sources are generally important, while English Wikipedia can be beneficial when non-English sources provide insufficient information for query generation. Based on these findings, we release four ToT test collections with 5,000 queries per language across multiple domains. This work provides the first large-scale multilingual ToT benchmark and offers practical guidance for constructing realistic ToT datasets beyond English.
Xuhong He, To Eun Kim, Maik Fröbe, Jaime Arguello, Bhaskar Mitra 0001, Fernando Diaz 0001
SIGIR3
2026 Formalized Information Needs Improve Large-Language-Model Relevance Judgments
abstract
Cranfield-style retrieval evaluations with too few or too many relevant documents or with low inter-assessor agreement on relevance can reduce the reliability of observations. In evaluations with human assessors, information needs are often formalized as retrieval topics to avoid an excessive number of relevant documents while maintaining good agreement. However, emerging evaluation setups that use Large Language Models (LLMs) as relevance assessors often use only queries, potentially decreasing the reliability. To study whether LLM relevance assessors benefit from formalized information needs, we synthetically formalize information needs with LLMs into topics that follow the established structure from previous human relevance assessments (i.e., descriptions and narratives). We compare assessors using synthetically formalized topics against the LLM-default query-only assessor on the~2019/2020~editions of TREC Deep Learning and Robust04. We find that assessors without formalization judge many more documents relevant and have a lower agreement, leading to reduced reliability in retrieval evaluations. Furthermore, we show that the formalized topics improve agreement between human and LLM relevance judgments, even when the topics are not highly similar to their human counterparts. Our findings indicate that LLM relevance assessors should use formalized information needs, as is standard for human assessment, and synthetically formalize topics when no human formalization exists to improve evaluation reliability.
Jüri Keller, Maik Fröbe, Björn Engelmann 0002, Fabian Haak 0001, Timo Breuer 0002, Birger Larsen, Philipp Schaer
SIGIR2
2025 Overview of PAN 2025: Generative AI Detection, Multilingual Text Detoxification, Multi-author Writing Style Analysis, and Generative Plagiarism Detection - Extended Abstract
Janek Bevendorff, Daryna Dementieva, Maik Fröbe, Bela Gipp, André Greiner-Petter, Jussi Karlgren, Maximilian Mayerl, Preslav Nakov, Alexander Panchenko, Martin Potthast, Artem Shelmanov, Efstathios Stamatatos, Benno Stein 0001, Yuxia Wang 0003, Matti Wiegmann, Eva Zangerle
ECIR (5)3
2025 The Second International Workshop on Open Web Search (WOWS)
Sheikh Mastura Farzana, Maik Fröbe, Michael Granitzer, Gijs Hendriksen, Djoerd Hiemstra, Martin Potthast, Arjen P. de Vries, Saber Zerhoudi
ECIR (5)2
2025 How Child-Friendly is Web Search? An Evaluation of Relevance vs. Harm
Maik Fröbe, Sophie Charlotte Bartholly, Matthias Hagen
ECIR (4)1
2025 Corpus Subsampling: Estimating the Effectiveness of Neural Retrieval Models on Large Corpora
Maik Fröbe, Andrew Parry, Harrisen Scells, Shuai Wang 0032, Shengyao Zhuang, Guido Zuccon, Martin Potthast, Matthias Hagen
ECIR (1)1
2025 Counterfactual Query Rewriting to Use Historical Relevance Feedback
Jüri Keller, Maik Fröbe, Gijs Hendriksen, Daria Alexander, Martin Potthast, Matthias Hagen, Philipp Schaer
ECIR (3)2
2025 Overview of Touché 2025: Argumentation Systems - Extended Abstract
Johannes Kiesel, Çagri Çöltekin, Marcel Gohsen, Sebastian Heineking, Maximilian Heinrich, Maik Fröbe, Tim Hagen, Mohammad Aliannejadi, Tomaz Erjavec, Matthias Hagen, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Harrisen Scells, Ines Zelch, Martin Potthast, Benno Stein 0001
ECIR (5)6
2025 Web-Scale Retrieval Experimentation with chatnoir-pyterrier
Jan Heinrich Merker, Janek Bevendorff, Maik Fröbe, Tim Hagen, Harrisen Scells, Matti Wiegmann, Benno Stein 0001, Matthias Hagen, Martin Potthast
ECIR (5)3
2025 Set-Encoder: Permutation-Invariant Inter-passage Attention for Listwise Passage Re-ranking with Cross-Encoders
Ferdinand Schlatt, Maik Fröbe, Harrisen Scells, Shengyao Zhuang, Bevan Koopman, Guido Zuccon, Benno Stein 0001, Martin Potthast, Matthias Hagen
ECIR (2)2
2025 Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-ranking
Ferdinand Schlatt, Maik Fröbe, Harrisen Scells, Shengyao Zhuang, Bevan Koopman, Guido Zuccon, Benno Stein 0001, Martin Potthast, Matthias Hagen
ECIR (3)2
2025 ReNeuIR at SIGIR 2025: The Fourth Workshop on Reaching Efficiency in Neural Information Retrieval
abstract
Measuring effectiveness and efficiency in information retrieval has a strong empirical background. While modern retrieval systems substantially improve effectiveness, the community has not yet agreed on how to measure efficiency, making it difficult to contrast effectiveness and efficiency fairly. Efficiency-oriented system comparisons are difficult due to factors such as hardware configurations, software versioning, and experimental settings. Efficiency affects users, researchers, and the environment and can be measured in many dimensions beyond time and space, such as resource consumption, water usage, and sample efficiency. Analyzing the efficiency of algorithms and their trade-off with effectiveness requires revisiting and establishing new standards and principles, from defining relevant concepts to designing new measures and guidelines to assess the findings' significance. ReNeuIR's fourth iteration aims to bring the community together to debate these questions and collaboratively test and improve benchmarking frameworks for efficiency based on discussions and collaborations of its previous iterations, including a shared task focused on efficiency and reproducibility.
Sebastian Bruch 0001, Maik Fröbe, Tim Hagen, Franco Maria Nardini, Martin Potthast
SIGIR2
2025 Large Language Model Relevance Assessors Agree With One Another More Than With Human Assessors
abstract
Relevance judgments can differ between assessors, but previous work has shown that such disagreements have little impact on the effectiveness rankings of retrieval systems. This applies to disagreements between humans as well as between human and large language model (LLM) assessors. However, the agreement between different LLM~assessors has not yet been systematically investigated. To close this gap, we compare eight LLM~assessors on the TREC DL tracks and the retrieval task of the RAG track with each other and with human assessors. We find that the agreement between LLM~assessors is higher than between LLMs and humans and, importantly, that LLM~assessors favor retrieval systems that use LLMs in their ranking decisions: our analyses with 30-50 retrieval systems show that the system rankings obtained by LLM~assessors overestimate LLM-based re-rankers by 9~to 17~positions on average.
Maik Fröbe, Andrew Parry, Ferdinand Schlatt, Sean MacAvaney, Benno Stein 0001, Martin Potthast, Matthias Hagen
SIGIR1
2025 The Viability of Crowdsourcing for RAG Evaluation
abstract
How good are humans at writing and judging responses in retrieval-augmented generation (RAG) scenarios? To answer this question, we investigate the efficacy of crowdsourcing for RAG through two complementary studies: response writing and response utility judgment. Our new Webis Crowd RAG Corpus 2025 (Webis-CrowdRAG-25) consists of 903 human-written and 903 LLM-generated responses for the 301 topics of the TREC 2024 RAG~track, with each response composed according to one of the three discourse styles 'bullet list', 'essay', or 'news'. For a selection of 65 topics, the corpus further contains 47,320 pairwise human judgments and 10,556 pairwise LLM judgments across seven utility dimensions (e.g., coverage and coherence). Our analyses give insights into human writing behavior for RAG and the viability of crowdsourcing for RAG evaluation. We find that human pairwise judgments provide reliable and cost-effective results. This is much less the case for LLM-based pairwise and human/LLM-based pointwise judgments, nor for automated comparisons with human-written reference responses. All our data and tools are freely available.
Lukas Gienapp, Tim Hagen, Maik Fröbe, Matthias Hagen, Benno Stein 0001, Martin Potthast, Harrisen Scells
SIGIR3
2025 TIREx Tracker: The Information Retrieval Experiment Tracker
abstract
The reproducibility and transparency of retrieval experiments depends on the availability of information about the experimental setup. However, the manual collection of experiment metadata can be tedious, error-prone, and inconsistent, which calls for an automated systematic collection. Expanding ir_metadata, we present the TIREx tracker, a tool that records hardware configurations, power/CPU/RAM/GPU usage, and experiment/system versions. Implemented as a lightweight platform-independent C binary, the TIREx tracker integrates seamlessly into Python, Java, or C/C++ workflows and can be easily integrated into shard task submissions, as we demonstrate for the TIRA/TIREx platform. Code, binaries, and documentation of the TIREx tracker are publicly available at https://github.com/tira-io/tirex-tracker.
Tim Hagen, Maik Fröbe, Jan Heinrich Merker, Harrisen Scells, Matthias Hagen, Martin Potthast
SIGIR2
2025 Variations in Relevance Judgments and the Shelf Life of Test Collections
abstract
The fundamental property of Cranfield-style evaluations, that system rankings are stable even when assessors disagree on individual relevance decisions, was validated on traditional test collections. However, the paradigm shift towards neural retrieval models affected the characteristics of modern test collections, e.g., documents are short, judged with four grades of relevance, and information needs have no descriptions or narratives. Under these changes, it is unclear whether assessor disagreement remains negligible for system comparisons. We investigate this aspect under the additional condition that the few modern test collections are heavily re-used. Given more possible query interpretations due to less formalized information needs, an ''expiration date'' for test collections might be needed if top-effectiveness requires overfitting to a single interpretation of relevance. We run a reproducibility study and re-annotate the relevance judgments of the 2019~TREC Deep Learning track. We can reproduce prior work in the neural retrieval setting, showing that assessor disagreement does not affect system rankings. However, we observe that some models substantially degrade with our new relevance judgments, and some have already reached the effectiveness of humans as rankers, providing evidence that test collections can expire.
Andrew Parry, Maik Fröbe, Harrisen Scells, Ferdinand Schlatt, Guglielmo Faggioli, Saber Zerhoudi, Sean MacAvaney, Eugene Yang 0001
SIGIR2
2025 Lightning IR: Straightforward Fine-tuning and Inference of Transformer-based Language Models for Information Retrieval
Ferdinand Schlatt, Maik Fröbe, Matthias Hagen
WSDM2
2024 The Eighth Workshop on Search-Oriented Conversational Artificial Intelligence (SCAI'24)
abstract
With the emergence of voice assistants and large language models, conversational interaction with information has become part of everyday life. The eighth edition of the search-oriented conversational AI (SCAI) workshop brings together practitioners and researchers from various disciplines to discuss challenges and advances in conversational search systems. This year’s edition focuses on evaluations beyond relevance and accuracy and looks at conversational search from the user’s perspective. The workshop features a shared task on user-centered evaluation datasets and metrics, challenging participants to develop new and innovative ways to evaluate conversational search systems while accounting for the needs and preferences of users.
Alexander Frummet, Andrea Papenmeier, Maik Fröbe, Johannes Kiesel
CHIIR3
2024 Overview of PAN 2024: Multi-author Writing Style Analysis, Multilingual Text Detoxification, Oppositional Thinking Analysis, and Generative AI Authorship Verification - Extended Abstract
Janek Bevendorff, Xavier Bonet Casals, Berta Chulvi, Daryna Dementieva, Ashraf Elnagar, Dayne Freitag, Maik Fröbe, Damir Korencic, Maximilian Mayerl, Animesh Mukherjee 0001, Alexander Panchenko, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Alisa Smirnova, Efstathios Stamatatos, Benno Stein 0001, Mariona Taulé, Dmitry Ustalov, Matti Wiegmann, Eva Zangerle
ECIR (6)7
2024 The First International Workshop on Open Web Search (WOWS)
Sheikh Mastura Farzana, Maik Fröbe, Michael Granitzer, Gijs Hendriksen, Djoerd Hiemstra, Martin Potthast, Saber Zerhoudi
ECIR (5)2
2024 The Open Web Index - Crawling and Indexing the Web for Public Use
Gijs Hendriksen, Michael Dinzinger, Sheikh Mastura Farzana, Noor Afshan Fathima, Maik Fröbe, Sebastian Heineking, Saber Zerhoudi, Michael Granitzer, Matthias Hagen, Djoerd Hiemstra, Martin Potthast, Benno Stein 0001
ECIR (5)5
2024 Overview of Touché 2024: Argumentation Systems
Johannes Kiesel, Çagri Çöltekin, Maximilian Heinrich, Maik Fröbe, Milad Alshomary, Bertrand De Longueville, Tomaz Erjavec, Nicolas Handke, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Theresa Reitis-Münstermann, Mario Scharfbillig, Nicolas Stefanovitch, Henning Wachsmuth, Martin Potthast, Benno Stein 0001
ECIR (5)4
2024 Analyzing Adversarial Attacks on Sequence-to-Sequence Relevance Models
Andrew Parry, Maik Fröbe, Sean MacAvaney, Martin Potthast, Matthias Hagen
ECIR (2)2
2024 Investigating the Effects of Sparse Attention on Cross-Encoders
Ferdinand Schlatt, Maik Fröbe, Matthias Hagen
ECIR (1)2
2024 ReNeuIR at SIGIR 2024: The Third Workshop on Reaching Efficiency in Neural Information Retrieval
Maik Fröbe, Joel Mackenzie, Bhaskar Mitra 0001, Franco Maria Nardini, Martin Potthast
SIGIR1
2024 Resources for Combining Teaching and Research in Information Retrieval Coursework
abstract
The first International Workshop on Open Web Search (WOWS) was held on Thursday, March 28th, at ECIR 2024 in Glasgow, UK. The full-day workshop had two calls for contributions: the first call aimed at scientific contributions to building, operating, and evaluating search engines cooperatively and the cooperative use of the web as a resource for researchers and innovators. The second call for implementations of retrieval components aimed to gain practical experience with joint, cooperative evaluation of search engines and their components. In total, 2~papers were accepted for the first call, and 11~software components were submitted for the second. The workshop ended with breakout sessions on how the OpenWebSearch.eu project can incorporate collaborative evaluations and a hub of search engines.
Maik Fröbe, Harrisen Scells, Theresa Elstner, Christopher Akiki, Lukas Gienapp, Jan Heinrich Merker, Sean MacAvaney, Benno Stein 0001, Matthias Hagen, Martin Potthast
SIGIR1
2024 Evaluating Generative Ad Hoc Information Retrieval
abstract
Recent advances in large language models have enabled the development of viable generative retrieval systems. Instead of a traditional document ranking, generative retrieval systems often directly return a grounded generated text as a response to a query. Quantifying the utility of the textual responses is essential for appropriately evaluating such generative ad hoc retrieval. Yet, the established evaluation methodology for ranking-based ad hoc retrieval is not suited for the reliable and reproducible evaluation of generated responses. To lay a foundation for developing new evaluation methods for generative retrieval systems, we survey the relevant literature from the fields of information retrieval and natural language processing, identify search tasks and system architectures in generative retrieval, develop a new user model, and study its operationalization.
Lukas Gienapp, Harrisen Scells, Niklas Deckers, Janek Bevendorff, Shuai Wang 0032, Johannes Kiesel, Shahbaz Syed, Maik Fröbe, Guido Zuccon, Benno Stein 0001, Matthias Hagen, Martin Potthast
SIGIR8
2024 Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
abstract
The zero-shot effectiveness of neural retrieval models is often evaluated on the BEIR benchmark---a combination of different IR evaluation datasets. Interestingly, previous studies found that particularly on the BEIR~subset Touché 2020, an argument retrieval task, neural retrieval models are considerably less effective than BM25. Still, so far, no further investigation has been conducted on what makes argument retrieval so "special''. To more deeply analyze the respective potential limits of neural retrieval models, we run a reproducibility study on the Touché 2020 data. In our study, we focus on two experiments: (i) a black-box evaluation (i.e., no model retraining), incorporating a theoretical exploration using retrieval axioms, and (ii) a data denoising evaluation involving post-hoc relevance judgments. Our black-box evaluation reveals an inherent bias of neural models towards retrieving short passages from the Touché 2020 data, and we also find that quite a few of the neural models' results are unjudged in the Touché 2020 data. As many of the short Touché passages are not argumentative and thus non-relevant per se, and as the missing judgments complicate fair comparison, we denoise the Touché 2020 data by excluding very short passages (less than 20 words) and by augmenting the unjudged data with post-hoc judgments following the Touché guidelines. On the denoised data, the effectiveness of the neural models improves by up to 0.52 in nDCG@10, but BM25 is still more effective. Our code and the augmented Touché 2020 dataset are available at https://github.com/castorini/touche-error-analysis.
Nandan Thakur, Luiz Bonifacio, Maik Fröbe, Alexander Bondarenko 0001, Ehsan Kamalloo, Martin Potthast, Matthias Hagen, Jimmy Lin
SIGIR3
2023 The Infinite Index: Information Retrieval on Generative Text-To-Image Models
abstract
Conditional generative models such as DALL-E and Stable Diffusion generate images based on a user-defined text, the prompt. Finding and refining prompts that produce a desired image has become the art of prompt engineering. Generative models do not provide a built-in retrieval model for a user’s information need expressed through prompts. In light of an extensive literature review, we reframe prompt engineering for generative models as interactive text-based retrieval on a novel kind of “infinite index”. We apply these insights for the first time in a case study on image generation for game design with an expert. Finally, we envision how active learning may help to guide the retrieval of generated images.
Niklas Deckers, Maik Fröbe, Johannes Kiesel, Gianluca Pandolfo, Christopher Schröder 0001, Benno Stein 0001, Martin Potthast
CHIIR2
2023 Overview of Touché 2023: Argument and Causal Retrieval - Extended Abstract
Alexander Bondarenko 0001, Maik Fröbe, Johannes Kiesel, Ferdinand Schlatt, Valentin Barrière, Brian Ravenet, Léo Hemamou, Simon Luck, Jan Heinrich Merker, Benno Stein 0001, Martin Potthast, Matthias Hagen
ECIR (3)2
2023 Bootstrapped nDCG Estimation in the Presence of Unjudged Documents
Maik Fröbe, Lukas Gienapp, Martin Potthast, Matthias Hagen
ECIR (1)1
2023 Continuous Integration for Reproducible Shared Tasks with TIRA.io
Maik Fröbe, Matti Wiegmann, Nikolay Kolyada, Bastian Grahm, Theresa Elstner, Frank Loebe, Matthias Hagen, Benno Stein 0001, Martin Potthast
ECIR (3)1
2023 On Stance Detection in Image Retrieval for Argumentation
abstract
Given a text query on a controversial topic, the task of Image Retrieval for Argumentation is to rank images according to how well they can be used to support a discussion on the topic. An important subtask therein is to determine the stance of the retrieved images, i.e., whether an image supports the pro or con side of the topic. In this paper, we conduct a comprehensive reproducibility study of the state of the art as represented by the CLEF'22 Touché lab and an in-house extension of it. Based on the submitted approaches, we developed a unified and modular retrieval process and reimplemented the submitted approaches according to this process. Through this unified reproduction (which also includes models not previously considered), we achieve an effectiveness improvement in argumentative image detection of up to 0.832 [email protected] However, despite this reproduction success, our study also revealed a previously unknown negative result: for stance detection, none of the reproduced or new approaches can convincingly beat a random baseline. To understand the apparent challenges inherent to image stance detection, we conduct a thorough error analysis and provide insight into potential new ways to approach this task.
Miriam Louise Carnot, Lorenz Heinemann, Jan Braker, Tobias Schreieder, Johannes Kiesel, Maik Fröbe, Martin Potthast, Benno Stein 0001
SIGIR6
2023 The Information Retrieval Experiment Platform
abstract
We integrate irdatasets, ir_measures, and PyTerrier with TIRA in the Information Retrieval Experiment Platform (TIREx) to promote more standardized, reproducible, scalable, and even blinded retrieval experiments. Standardization is achieved when a retrieval approach implements PyTerrier's interfaces and the input and output of an experiment are compatible with ir_datasets and ir_measures. However, none of this is a must for reproducibility and scalability, as TIRA can run any dockerized software locally or remotely in a cloud-native execution environment. Version control and caching ensure efficient (re)execution. TIRA allows for blind evaluation when an experiment runs on a remote server or cloud not under the control of the experimenter. The test data and ground truth are then hidden from public access, and the retrieval software has to process them in a sandbox that prevents data leaks.
Maik Fröbe, Jan Heinrich Merker, Sean MacAvaney, Niklas Deckers, Simon Reich, Janek Bevendorff, Benno Stein 0001, Matthias Hagen, Martin Potthast
SIGIR1
2023 The Archive Query Log: Mining Millions of Search Result Pages of Hundreds of Search Engines from 25 Years of Web Archives
abstract
The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years. Its first version includes 356 million queries, 137 million search result pages, and 1.4 billion search results across 550 search providers. Although many query logs have been studied in the literature, the search providers that own them generally do not publish their logs to protect user privacy and vital business data. Of the few query logs publicly available, none combines size, scope, and diversity. The AQL is the first to do so, enabling research on new retrieval models and (diachronic) search engine analyses. Provided in a privacy-preserving manner, it promotes open research as well as more transparency and accountability in the search industry.
Jan Heinrich Merker, Sebastian Heineking, Maik Fröbe, Lukas Gienapp, Harrisen Scells, Benno Stein 0001, Matthias Hagen, Martin Potthast
SIGIR3
2022 Overview of Touché 2022: Argument Retrieval - Extended Abstract
Alexander Bondarenko 0001, Maik Fröbe, Johannes Kiesel, Shahbaz Syed, Timon Ziegenbein, Meriem Beloucif, Alexander Panchenko, Chris Biemann, Benno Stein 0001, Henning Wachsmuth, Martin Potthast, Matthias Hagen
ECIR (2)2
2022 The Power of Anchor Text in the Neural Retrieval Era
Maik Fröbe, Sebastian Günther 0002, Maximilian Probst Gutenberg, Martin Potthast, Matthias Hagen
ECIR (1)1
2022 City of Disguise: A Query Obfuscation Game on the ClueWeb
Maik Fröbe, Nicola Lea Libera, Matthias Hagen
ECIR (2)1
2022 Axiomatic Retrieval Experimentation with ir_axioms
abstract
Axiomatic approaches to information retrieval have played a key role in determining basic constraints that characterize good retrieval models. Beyond their importance in retrieval theory, axioms have been operationalized to improve an initial ranking, to "guide" retrieval, or to explain some model's rankings. However, recent open-source retrieval frameworks like PyTerrier and Pyserini, which made it easy to experiment with sparse and dense retrieval models, have not included any retrieval axiom support so far.
Alexander Bondarenko 0001, Maik Fröbe, Jan Heinrich Merker, Benno Stein 0001, Michael Völske, Matthias Hagen
SIGIR2
2022 How Train-Test Leakage Affects Zero-Shot Retrieval
Maik Fröbe, Christopher Akiki, Martin Potthast, Matthias Hagen
SPIRE1
2021 Overview of Touché 2021: Argument Retrieval - Extended Abstract
Alexander Bondarenko 0001, Lukas Gienapp, Maik Fröbe, Meriem Beloucif, Yamen Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein 0001, Henning Wachsmuth, Martin Potthast, Matthias Hagen
ECIR (2)3
2021 CopyCat: Near-Duplicates Within and Between the ClueWeb and the Common Crawl
abstract
The amount of near-duplicates in web crawls like the ClueWeb or Common Crawl demands from their users either to develop a preprocessing pipeline for deduplication, which is costly both computationally and in person hours, or accepting the undesired effects that near-duplicates have on reliability and validity of experiments. We introduce ChatNoir-CopyCat-21, which simplifies deduplication significantly. It comes in two parts: (1) A compilation of near-duplicate documents within the ClueWeb09, the ClueWeb12, and two Common Crawl snapshots, as well as between selections of these crawls, and (2) a software library that implements the deduplication of arbitrary document sets. Our analysis shows that 14--52, of the documents within a crawl and around~0.7--2.5, between the crawls are near-duplicates. Two showcases demonstrate the application and usefulness of our resource.
Maik Fröbe, Janek Bevendorff, Lukas Gienapp, Michael Völske, Benno Stein 0001, Martin Potthast, Matthias Hagen
SIGIR1
2021 The Information Retrieval Anthology
abstract
We present the IR Anthology, a corpus of information retrieval publications accessible via a metadata browser and a full-text search engine. Following the example of the well-known ACL Anthology, the IR Anthology serves as a hub for researchers interested in information retrieval. Our search engine ChatNoir indexes the publications' full texts, enabling a focused search and linking users to the respective publisher's site for personal access. Listing more than 40,000 publications at the time of writing, the IR Anthology can be freely accessed at https://IR.webis.de.
Martin Potthast, Sebastian Günther 0002, Janek Bevendorff, Jan Philipp Bittner, Alexander Bondarenko 0001, Maik Fröbe, Christian Kahmann, Andreas Niekler, Michael Völske, Benno Stein 0001, Matthias Hagen
SIGIR6
2020 The Impact of Negative Relevance Judgments on NDCG
abstract
NDCG is one of the most commonly used measures to quantify system performance in retrieval experiments. Though originally not considered, graded relevance judgments nowadays frequently include negative labels. Negative relevance labels cause NDCG to be unbounded. This is probably why widely used implementations of NDCG map negative relevance labels to zero, thus ensuring the resulting scores to originate from the [0,1] range. But zeroing negative labels discards valuable relevance information, e.g., by treating spam documents the same as unjudged ones, which are assigned the relevance label of zero by default. We show that, instead of zeroing negative labels, a min-max-normalization of NDCG retains its statistical power while improving its reliability and stability.
Lukas Gienapp, Maik Fröbe, Matthias Hagen, Martin Potthast
CIKM2
2020 The Effect of Content-Equivalent Near-Duplicates on the Evaluation of Search Engines
Maik Fröbe, Jan Philipp Bittner, Martin Potthast, Matthias Hagen
ECIR (2)1
2020 A Search Engine for Police Press Releases to Double-Check the News
Maik Fröbe, Nina Schwanke, Matthias Hagen, Martin Potthast
ECIR (2)1
2020 Sampling Bias Due to Near-Duplicates in Learning to Rank
abstract
Learning to rank~(LTR) is the de facto standard for web search, improving upon classical retrieval models by exploiting (in)direct relevance feedback from user judgments, interaction logs, etc. We investigate for the first time the effect of a sampling bias on LTR~models due to the potential presence of near-duplicate web pages in the training data, and how (in)consistent relevance feedback of duplicates influences an LTR~model's decisions. To examine this bias, we construct a series of specialized LTR~datasets based on the ClueWeb09 corpus with varying amounts of near-duplicates. We devise worst-case and average-case train/test splits that are evaluated on popular pointwise, pairwise, and listwise LTR~models. Our experiments demonstrate that duplication causes overfitting and thus less effective models, making a strong case for the benefits of systematic deduplication before training and model evaluation.
Maik Fröbe, Janek Bevendorff, Jan Heinrich Merker, Martin Potthast, Matthias Hagen
SIGIR1
2020 Comparative Web Search Questions
abstract
\beginabstract We analyze comparative questions, i.e., questions asking to compare different items, that were submitted to Yandex in 2012. Responses to such questions might be quite different from the simple "ten blue links'' and could, for example, aggregate pros and cons of the different options as direct answers. However, changing the result presentation is an intricate decision such that the classification of comparative questions forms a highly precision-oriented task.
Alexander Bondarenko 0001, Pavel Braslavski 0001, Michael Völske, Rami Aly, Maik Fröbe, Alexander Panchenko, Chris Biemann, Benno Stein 0001, Matthias Hagen
WSDM5