Martin Potthast

dblp:87/6573 · DBLP profile ↗
← Back
112ranked-venue papers in the field
15as first author
67since 2021 · last 2026
0000-0003-2451-0665ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 109 (14 first)Data Mining & Knowledge Discovery · 3 (1 first)
YearPublicationVenuePosition
2026 Overview of PAN 2026: Voight-Kampff Generative AI Detection, Text Watermarking, Multi-author Writing Style Analysis, Generative Plagiarism Detection, and Reasoning Trajectory Detection
Janek Bevendorff, Maik Fröbe, André Greiner-Petter, Andreas Jakoby, Maximilian Mayerl, Preslav Nakov, Henry Plutz, Martin Potthast, Benno Stein 0001, Minh Ngoc Ta, Yuxia Wang 0003, Eva Zangerle
ECIR (4)8
2026 The Third International Workshop on Open Web Search (WOWS)
Laura Caspari, Maik Fröbe, Sebastian Heineking, Michael Granitzer, Gijs Hendriksen, Djoerd Hiemstra, Martin Potthast, Arjen P. de Vries, Saber Zerhoudi
ECIR (3)7
2026 Evaluating the Efficiency and Effectiveness of Learned Sparse Retrieval with the lsr_benchmark
Maik Fröbe, Ferdinand Schlatt, Cosimo Rulli, Tim Hagen, Jan Heinrich Merker, Gijs Hendriksen, Carlos Eduardo Rosar Kós Lassance, Franco Maria Nardini, Rossano Venturini, Martin Potthast
ECIR (4)10
2026 Overview of Touché 2026: Argumentation Systems - Extended Abstract
Johannes Kiesel, Marc Feger, Tim Hagen, Sebastian Heineking, Maximilian Heinrich, Maik Fröbe, Katarina Boland, Wilhelm Pertsch, Julia Romberg, Ines Zelch, Stefan Dietze, Matthias Hagen, Martin Potthast, Benno Stein 0001
ECIR (4)13
2026 An Open SERP Mining Infrastructure for the Archive Query Log
Jan Heinrich Merker, Simon Ruth, Harrisen Scells, Martin Potthast
ECIR (4)4
2026 Creating Specialized RAG-Based Search Engines Using the Open Web Index
Alexander Nussbaumer, Michael Dinzinger, Sebastian Heineking, Gijs Hendriksen, Felix Holz, Saber Zerhoudi, Martin Potthast, Michael Granitzer
ECIR (4)7
2026 Auto-Judge: A Cross-Task Benchmark for Comparing LLM Judges for Citation-Grounded RAG Systems
abstract
We present the Auto-Judge resource for the meta-evaluation of automated LLM judges, especially judges that evaluate Retrieval-Augmented Generation (RAG) systems that ground their response with citations. The resource couples (i) a data release of topics, pooled RAG responses, and human judgments, with (ii) a standardized protocol and software infrastructure for implementing "LLM-as-a-judge" methods in a reproducible and extensible way, including support for parameter sweeps and variant tracking.
Naghmeh Farzi, Tim Hagen, Eugene Yang 0001, Maik Fröbe, Ronak Pradeep, Hossein A. Rahmani, Xi Wang 0012, Oleg Zendel, Martin Potthast, Laura Dietz
SIGIR9
2026 ReNeuIR at SIGIR 2026: The Fifth Workshop on Reaching Efficiency in Neural Information Retrieval
abstract
The lack of efficiency in neural information retrieval remains one of the primary obstacles to deploying neural retrieval models as a first-stage retriever at scale. While recent tools have improved the standardized measurement of model efficiency, substantial progress is still needed to enable systematic comparative evaluation, for example, in terms of standards for systems and hardware configurations, cloud-based evaluation, benchmarks, and reproducibility. Beyond measurement, the IR community also needs stronger incentives to move in this direction, such as cost-efficiency as a review criterion or as efficiency and/or effectiveness measures in shared tasks, related teaching materials, efficiency-oriented user studies, and specialized awards for efficiency achievements. In particular, developing more efficient variants of highly effective retrieval algorithms should become an admissible research goal for PhD students if cost-efficiency is to become a first-class design objective in~IR. With ReNeuIR, we have established a recurring forum where these questions and new ideas are discussed and where the community comes together to collaboratively evaluate and improve efficiency benchmarking frameworks---most notably through the organization of a shared task focused on efficiency and reproducibility.
Maik Fröbe, Tim Hagen, Franco Maria Nardini, Martin Potthast
SIGIR4
2026 Humans, LLMs, and Measures Do Not Align in Attributed Information Retrieval
abstract
Evaluating attributed information retrieval (AIR) systems requires assessing both informativeness and attributability. To enable scalable evaluation, LLM-sourced ground truth data is frequently used, yet the validity of this practice remains unclear. We replicate the evaluation framework of Djeddal et al. [1] which relies on LLM-written ground truth answers, and additionally crowdsource human-written answers and pairwise preference judgments. This allows us to investigate (1) how robust reference-based evaluation measures are to gold reference variation; (2) to what extent do LLM judges agree with human annotators; and (3) which automatic measures best predict human and LLM preferences? Our findings reveal substantial sensitivity of reference-based measures to gold reference choice, and human and LLM judges exhibiting low agreement on preference judgments, despite similar aggregate tendencies. Furthermore, no automatic evaluation measure strongly predicts human preferences, suggesting a fundamental methodological shortcoming in current AIR evaluation practices.
Lukas Gienapp, Jenny Lang, Martin Potthast, Harrisen Scells
SIGIR3
2026 Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
abstract
The unjudged document problem, where systems that did not contribute to the original judgement pool may retrieve documents without a relevance judgement, is a key obstacle to the reuseability of test collections in information retrieval. While the de facto standard to deal with the problem is to treat unjudged documents as non-relevant, many alternatives have been proposed, such as the use of large language models (LLMs) as a relevance judge (LLM-as-a-judge). However, this has been criticized, among other things, as circular, since the same LLM can be used as the ranker and the judge. We propose to train topic-specific relevance classifiers instead: By finetuning monoT5 with independent LoRA weight adaptation on the judgments of a single assessor for a single topic's pool, we align it to that assessor's notion of relevance for the topic. The system rankings obtained through our classifier's relevance judgments achieve a Spearmans' $ρ$ correlation of $>0.94$ with ground truth system rankings. As little as 128 initial human judgments per topic suffice to improve the comparability of models, compared to treating unjudged documents as non-relevant, while achieving more reliability than existing LLM-as-a-judge approaches. Topic-specific relevance classifiers are thus a lightweight and straightforward way to tackle the unjudged document problem, while maintaining human judgments as the gold standard for retrieval evaluation. Code, models, and data are made openly available.
Lukas Gienapp, Martin Potthast, Andrew Yates, Harrisen Scells, Eugene Yang 0001
SIGIR2
2025 Overview of PAN 2025: Generative AI Detection, Multilingual Text Detoxification, Multi-author Writing Style Analysis, and Generative Plagiarism Detection - Extended Abstract
Janek Bevendorff, Daryna Dementieva, Maik Fröbe, Bela Gipp, André Greiner-Petter, Jussi Karlgren, Maximilian Mayerl, Preslav Nakov, Alexander Panchenko, Martin Potthast, Artem Shelmanov, Efstathios Stamatatos, Benno Stein 0001, Yuxia Wang 0003, Matti Wiegmann, Eva Zangerle
ECIR (5)10
2025 The Second International Workshop on Open Web Search (WOWS)
Sheikh Mastura Farzana, Maik Fröbe, Michael Granitzer, Gijs Hendriksen, Djoerd Hiemstra, Martin Potthast, Arjen P. de Vries, Saber Zerhoudi
ECIR (5)6
2025 Corpus Subsampling: Estimating the Effectiveness of Neural Retrieval Models on Large Corpora
Maik Fröbe, Andrew Parry, Harrisen Scells, Shuai Wang 0032, Shengyao Zhuang, Guido Zuccon, Martin Potthast, Matthias Hagen
ECIR (1)7
2025 Call for Research on the Impact of Information Retrieval on Social Norms
Tim Gollub, Pierre Achkar, Martin Potthast, Benno Stein 0001
ECIR (4)3
2025 Ranking Generated Answers - On the Agreement of Retrieval Models with Humans on Consumer Health Questions
Sebastian Heineking, Jonas Probst, Daniel Steinbach, Martin Potthast, Harrisen Scells
ECIR (3)4
2025 ImageCLEF 2025: Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications
Bogdan Ionescu, Henning Müller, Dan-Cristian Stanciu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Andrea M. Storås, Asma Ben Abacha, Benjamin Bracke, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Diandra Fabre, Didier Schwab, Dimitar Dimitrov 0003, Emmanuelle Esperança-Rodier, Mihai Gabriel Constantin, Helmut Becker, Hendrik Damm, Henning Schäfer, Ivan Rodkin, Ivan Koychev, Johannes Kiesel, Johannes Rückert, Josep Malvehy, Liviu-Daniel Stefan, Louise Bloch, Martin Potthast, Maximilian Heinrich, Michael Riegler 0001, Mihai Dogariu, Noel Codella, Pål Halvorsen, Preslav Nakov, Raphael Brüngel, Roberto A. Novoa, Rocktim Jyoti Das, Steven Alexander Hicks, Sushant Gautam, Tabea Margareta Grace Pakull, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim, Zhuohan Xie
ECIR (5)31
2025 Counterfactual Query Rewriting to Use Historical Relevance Feedback
Jüri Keller, Maik Fröbe, Gijs Hendriksen, Daria Alexander, Martin Potthast, Matthias Hagen, Philipp Schaer
ECIR (3)5
2025 Overview of Touché 2025: Argumentation Systems - Extended Abstract
Johannes Kiesel, Çagri Çöltekin, Marcel Gohsen, Sebastian Heineking, Maximilian Heinrich, Maik Fröbe, Tim Hagen, Mohammad Aliannejadi, Tomaz Erjavec, Matthias Hagen, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Harrisen Scells, Ines Zelch, Martin Potthast, Benno Stein 0001
ECIR (5)18
2025 A Test Collection for Dataset Retrieval
Nikolay Kolyada, Martin Potthast, Benno Stein 0001
ECIR (3)2
2025 Web-Scale Retrieval Experimentation with chatnoir-pyterrier
Jan Heinrich Merker, Janek Bevendorff, Maik Fröbe, Tim Hagen, Harrisen Scells, Matti Wiegmann, Benno Stein 0001, Matthias Hagen, Martin Potthast
ECIR (5)9
2025 Set-Encoder: Permutation-Invariant Inter-passage Attention for Listwise Passage Re-ranking with Cross-Encoders
Ferdinand Schlatt, Maik Fröbe, Harrisen Scells, Shengyao Zhuang, Bevan Koopman, Guido Zuccon, Benno Stein 0001, Martin Potthast, Matthias Hagen
ECIR (2)8
2025 Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-ranking
Ferdinand Schlatt, Maik Fröbe, Harrisen Scells, Shengyao Zhuang, Bevan Koopman, Guido Zuccon, Benno Stein 0001, Martin Potthast, Matthias Hagen
ECIR (3)8
2025 ReNeuIR at SIGIR 2025: The Fourth Workshop on Reaching Efficiency in Neural Information Retrieval
abstract
Measuring effectiveness and efficiency in information retrieval has a strong empirical background. While modern retrieval systems substantially improve effectiveness, the community has not yet agreed on how to measure efficiency, making it difficult to contrast effectiveness and efficiency fairly. Efficiency-oriented system comparisons are difficult due to factors such as hardware configurations, software versioning, and experimental settings. Efficiency affects users, researchers, and the environment and can be measured in many dimensions beyond time and space, such as resource consumption, water usage, and sample efficiency. Analyzing the efficiency of algorithms and their trade-off with effectiveness requires revisiting and establishing new standards and principles, from defining relevant concepts to designing new measures and guidelines to assess the findings' significance. ReNeuIR's fourth iteration aims to bring the community together to debate these questions and collaboratively test and improve benchmarking frameworks for efficiency based on discussions and collaborations of its previous iterations, including a shared task focused on efficiency and reproducibility.
Sebastian Bruch 0001, Maik Fröbe, Tim Hagen, Franco Maria Nardini, Martin Potthast
SIGIR5
2025 Large Language Model Relevance Assessors Agree With One Another More Than With Human Assessors
abstract
Relevance judgments can differ between assessors, but previous work has shown that such disagreements have little impact on the effectiveness rankings of retrieval systems. This applies to disagreements between humans as well as between human and large language model (LLM) assessors. However, the agreement between different LLM~assessors has not yet been systematically investigated. To close this gap, we compare eight LLM~assessors on the TREC DL tracks and the retrieval task of the RAG track with each other and with human assessors. We find that the agreement between LLM~assessors is higher than between LLMs and humans and, importantly, that LLM~assessors favor retrieval systems that use LLMs in their ranking decisions: our analyses with 30-50 retrieval systems show that the system rankings obtained by LLM~assessors overestimate LLM-based re-rankers by 9~to 17~positions on average.
Maik Fröbe, Andrew Parry, Ferdinand Schlatt, Sean MacAvaney, Benno Stein 0001, Martin Potthast, Matthias Hagen
SIGIR6
2025 The Viability of Crowdsourcing for RAG Evaluation
abstract
How good are humans at writing and judging responses in retrieval-augmented generation (RAG) scenarios? To answer this question, we investigate the efficacy of crowdsourcing for RAG through two complementary studies: response writing and response utility judgment. Our new Webis Crowd RAG Corpus 2025 (Webis-CrowdRAG-25) consists of 903 human-written and 903 LLM-generated responses for the 301 topics of the TREC 2024 RAG~track, with each response composed according to one of the three discourse styles 'bullet list', 'essay', or 'news'. For a selection of 65 topics, the corpus further contains 47,320 pairwise human judgments and 10,556 pairwise LLM judgments across seven utility dimensions (e.g., coverage and coherence). Our analyses give insights into human writing behavior for RAG and the viability of crowdsourcing for RAG evaluation. We find that human pairwise judgments provide reliable and cost-effective results. This is much less the case for LLM-based pairwise and human/LLM-based pointwise judgments, nor for automated comparisons with human-written reference responses. All our data and tools are freely available.
Lukas Gienapp, Tim Hagen, Maik Fröbe, Matthias Hagen, Benno Stein 0001, Martin Potthast, Harrisen Scells
SIGIR6
2025 TIREx Tracker: The Information Retrieval Experiment Tracker
abstract
The reproducibility and transparency of retrieval experiments depends on the availability of information about the experimental setup. However, the manual collection of experiment metadata can be tedious, error-prone, and inconsistent, which calls for an automated systematic collection. Expanding ir_metadata, we present the TIREx tracker, a tool that records hardware configurations, power/CPU/RAM/GPU usage, and experiment/system versions. Implemented as a lightweight platform-independent C binary, the TIREx tracker integrates seamlessly into Python, Java, or C/C++ workflows and can be easily integrated into shard task submissions, as we demonstrate for the TIRA/TIREx platform. Code, binaries, and documentation of the TIREx tracker are publicly available at https://github.com/tira-io/tirex-tracker.
Tim Hagen, Maik Fröbe, Jan Heinrich Merker, Harrisen Scells, Matthias Hagen, Martin Potthast
SIGIR6
2025 AiReview: An Open Platform for Accelerating Systematic Reviews with LLMs
abstract
Systematic reviews are fundamental to evidence-based medicine.Creating one is time-consuming and labour-intensive, mainly due to the need to screen, or assess, many studies for inclusion in the review.Existing tools help streamline this process, mostly using traditional machine learning.Large language models (LLMs) offer new opportunities to speed up screening, yet no tool currently enables users to directly apply LLMs or ensures systematic and transparent use of these methods.This paper presents (i) a flexible framework for using LLMs in systematic review tasks, especially title and abstract screening, and (ii) a web-based interface for LLMassisted screening.Together, they form AiReview-a novel platform that connects cutting-edge LLM-assisted screening methods with real-world systematic review practice.The live tool is available at https://aireview.ielab.io.We also release the code publicly at https://github.com/ielab/ai-review.
Xinyu Mao 0001, Teerapong Leelanupab, Martin Potthast, Harrisen Scells, Guido Zuccon
SIGIR3
2025 TITE: Token-Independent Text Encoder for Information Retrieval
abstract
Transformer-based retrieval approaches typically use the contextualized embedding of the first input token as a dense vector representation for queries and documents. The embeddings of all other tokens are also computed but then discarded, wasting resources. In this paper, we propose the Token-Independent Text Encoder (TITE) as a more efficient modification of the backbone encoder model. Using an attention-based pooling technique, TITE iteratively reduces the sequence length of hidden states layer by layer so that the final output is already a single sequence representation vector. Our empirical analyses on the TREC 2019 and 2020 Deep Learning tracks and the BEIR benchmark show that TITE is on par in terms of effectiveness compared to standard bi-encoder retrieval models while being up to 3.3 times faster at encoding queries and documents. Our code is available at: https://github.com/webis-de/SIGIR-25.
Ferdinand Schlatt, Tim Hagen, Martin Potthast, Matthias Hagen
SIGIR3
2024 Product Spam on YouTube: A Case Study
abstract
YouTube videos are a popular medium for online product reviews. They are not only informative and entertaining, but may also be perceived as quite credible under the viewer’s impression of a personal product demonstration by an expert. As the world’s largest online video platform, YouTube’s content is included prominently in the results of most general-purpose web search engines. Consequently, online marketeers are using classic Search Engine Optimization (SEO) techniques also for placing their video content in search engines. Over the years, we have noticed an ever increasing noise floor of low-quality SEO content in product search results and in this study, we show that this trend has spilled over into videos as well. We examine YouTube video reviews for several thousand products retrieved from three commercial search engines and conduct spam detection experiments based directly on the videos’ subtitle transcripts rather than relying on metadata and comments. We find that at least a third of the retrieved videos can be regarded as spam or low-quality productions. We are further able to distinguish these spam product reviews accurately from higher-quality videos with a semi-supervised n-gram classification approach.
Janek Bevendorff, Matti Wiegmann, Martin Potthast, Benno Stein 0001
CHIIR3
2024 A User Study on the Acceptance of Native Advertising in Generative IR
abstract
Commercial conversational search engines need a business model. Since advertising is the main source of revenue for “traditional” ten-blue-links web search, ads are not an unlikely option for conversational search either. In traditional web search, ads are usually placed above organic search results. However, large language models (LLMs) may be dynamically prompted to blend product placements with “organic” conversational responses, similar to native advertising in journalism. This type of advertising can be very difficult to recognize, depending on how subtly it is integrated and disclosed. To raise awareness of this potential development, we analyze the capabilities of current LLMs to blend ads with generative search results. In a user study, we ask people about the perceived quality of (emulated) search results in different advertising scenarios. In a substantial number of cases, our survey participants do not notice brand or product placements when they do not expect them. Thus, our results show the potential of LLMs to subtly mix advertising with generated search results. This warrants further investigation, for example, to develop appropriate advertising disclosure rules, and to detect advertising in generated results. Our research also raises broader concerns about whether commercial or open-source generative models can be trusted not to be fine-tuned to generate ads rather than “genuine” responses.
Ines Zelch, Matthias Hagen, Martin Potthast
CHIIR3
2024 Overview of PAN 2024: Multi-author Writing Style Analysis, Multilingual Text Detoxification, Oppositional Thinking Analysis, and Generative AI Authorship Verification - Extended Abstract
Janek Bevendorff, Xavier Bonet Casals, Berta Chulvi, Daryna Dementieva, Ashraf Elnagar, Dayne Freitag, Maik Fröbe, Damir Korencic, Maximilian Mayerl, Animesh Mukherjee 0001, Alexander Panchenko, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Alisa Smirnova, Efstathios Stamatatos, Benno Stein 0001, Mariona Taulé, Dmitry Ustalov, Matti Wiegmann, Eva Zangerle
ECIR (6)12
2024 Is Google Getting Worse? A Longitudinal Investigation of SEO Spam in Search Engines
Janek Bevendorff, Matti Wiegmann, Martin Potthast, Benno Stein 0001
ECIR (3)3
2024 The First International Workshop on Open Web Search (WOWS)
Sheikh Mastura Farzana, Maik Fröbe, Michael Granitzer, Gijs Hendriksen, Djoerd Hiemstra, Martin Potthast, Saber Zerhoudi
ECIR (5)6
2024 The Open Web Index - Crawling and Indexing the Web for Public Use
Gijs Hendriksen, Michael Dinzinger, Sheikh Mastura Farzana, Noor Afshan Fathima, Maik Fröbe, Sebastian Heineking, Saber Zerhoudi, Michael Granitzer, Matthias Hagen, Djoerd Hiemstra, Martin Potthast, Benno Stein 0001
ECIR (5)11
2024 Advancing Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications with ImageCLEF 2024
Bogdan Ionescu, Henning Müller, Ana-Maria Claudia Dragulinescu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandru Stan, Andrea M. Storås, Asma Ben Abacha, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Didier Schwab, Emmanuelle Esperança-Rodier, George Ioannidis, Griffin Adams, Henning Schäfer, Hugo Manguinhas, Ioan Coman, Johanna Schöler, Johannes Kiesel, Johannes Rückert, Louise Bloch, Martin Potthast, Maximilian Heinrich, Meliha Yetisgen, Michael Riegler 0001, Neal Snider, Pål Halvorsen, Raphael Brüngel, Steven Alexander Hicks, Vajira Thambawita, Vassili Kovalev, Yuri Prokopchuk, Wen-Wai Yim
ECIR (6)27
2024 Overview of Touché 2024: Argumentation Systems
Johannes Kiesel, Çagri Çöltekin, Maximilian Heinrich, Maik Fröbe, Milad Alshomary, Bertrand De Longueville, Tomaz Erjavec, Nicolas Handke, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Theresa Reitis-Münstermann, Mario Scharfbillig, Nicolas Stefanovitch, Henning Wachsmuth, Martin Potthast, Benno Stein 0001
ECIR (5)18
2024 Analyzing Adversarial Attacks on Sequence-to-Sequence Relevance Models
Andrew Parry, Maik Fröbe, Sean MacAvaney, Martin Potthast, Matthias Hagen
ECIR (2)4
2024 Zero-Shot Generative Large Language Models for Systematic Review Screening Automation
Shuai Wang 0032, Harrisen Scells, Shengyao Zhuang, Martin Potthast, Bevan Koopman, Guido Zuccon
ECIR (1)4
2024 ReNeuIR at SIGIR 2024: The Third Workshop on Reaching Efficiency in Neural Information Retrieval
Maik Fröbe, Joel Mackenzie, Bhaskar Mitra 0001, Franco Maria Nardini, Martin Potthast
SIGIR5
2024 Resources for Combining Teaching and Research in Information Retrieval Coursework
abstract
The first International Workshop on Open Web Search (WOWS) was held on Thursday, March 28th, at ECIR 2024 in Glasgow, UK. The full-day workshop had two calls for contributions: the first call aimed at scientific contributions to building, operating, and evaluating search engines cooperatively and the cooperative use of the web as a resource for researchers and innovators. The second call for implementations of retrieval components aimed to gain practical experience with joint, cooperative evaluation of search engines and their components. In total, 2~papers were accepted for the first call, and 11~software components were submitted for the second. The workshop ended with breakout sessions on how the OpenWebSearch.eu project can incorporate collaborative evaluations and a hub of search engines.
Maik Fröbe, Harrisen Scells, Theresa Elstner, Christopher Akiki, Lukas Gienapp, Jan Heinrich Merker, Sean MacAvaney, Benno Stein 0001, Matthias Hagen, Martin Potthast
SIGIR10
2024 Evaluating Generative Ad Hoc Information Retrieval
abstract
Recent advances in large language models have enabled the development of viable generative retrieval systems. Instead of a traditional document ranking, generative retrieval systems often directly return a grounded generated text as a response to a query. Quantifying the utility of the textual responses is essential for appropriately evaluating such generative ad hoc retrieval. Yet, the established evaluation methodology for ranking-based ad hoc retrieval is not suited for the reliable and reproducible evaluation of generated responses. To lay a foundation for developing new evaluation methods for generative retrieval systems, we survey the relevant literature from the fields of information retrieval and natural language processing, identify search tasks and system architectures in generative retrieval, develop a new user model, and study its operationalization.
Lukas Gienapp, Harrisen Scells, Niklas Deckers, Janek Bevendorff, Shuai Wang 0032, Johannes Kiesel, Shahbaz Syed, Maik Fröbe, Guido Zuccon, Benno Stein 0001, Matthias Hagen, Martin Potthast
SIGIR12
2024 Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
abstract
The zero-shot effectiveness of neural retrieval models is often evaluated on the BEIR benchmark---a combination of different IR evaluation datasets. Interestingly, previous studies found that particularly on the BEIR~subset Touché 2020, an argument retrieval task, neural retrieval models are considerably less effective than BM25. Still, so far, no further investigation has been conducted on what makes argument retrieval so "special''. To more deeply analyze the respective potential limits of neural retrieval models, we run a reproducibility study on the Touché 2020 data. In our study, we focus on two experiments: (i) a black-box evaluation (i.e., no model retraining), incorporating a theoretical exploration using retrieval axioms, and (ii) a data denoising evaluation involving post-hoc relevance judgments. Our black-box evaluation reveals an inherent bias of neural models towards retrieving short passages from the Touché 2020 data, and we also find that quite a few of the neural models' results are unjudged in the Touché 2020 data. As many of the short Touché passages are not argumentative and thus non-relevant per se, and as the missing judgments complicate fair comparison, we denoise the Touché 2020 data by excluding very short passages (less than 20 words) and by augmenting the unjudged data with post-hoc judgments following the Touché guidelines. On the denoised data, the effectiveness of the neural models improves by up to 0.52 in nDCG@10, but BM25 is still more effective. Our code and the augmented Touché 2020 dataset are available at https://github.com/castorini/touche-error-analysis.
Nandan Thakur, Luiz Bonifacio, Maik Fröbe, Alexander Bondarenko 0001, Ehsan Kamalloo, Martin Potthast, Matthias Hagen, Jimmy Lin
SIGIR6
2024 Impact and development of an Open Web Index for open web search
abstract
Abstract Web search is a crucial technology for the digital economy. Dominated by a few gatekeepers focused on commercial success, however, web publishers have to optimize their content for these gatekeepers, resulting in a closed ecosystem of search engines as well as the risk of publishers sacrificing quality. To encourage an open search ecosystem and offer users genuine choice among alternative search engines, we propose the development of an Open Web Index (OWI). We outline six core principles for developing and maintaining an open index, based on open data principles, legal compliance, and collaborative technology development. The combination of an open index with what we call declarative search engines will facilitate the development of vertical search engines and innovative web data products (including, e.g., large language models), enabling a fair and open information space. This framework underpins the EU‐funded project OpenWebSearch.EU, marking the first step towards realizing an Open Web Index.
Michael Granitzer, Stefan Voigt, Noor Afshan Fathima, Martin Golasowski, Christian Gütl, Tobias Hecking, Gijs Hendriksen, Djoerd Hiemstra, Jan Martinovic, Jelena Mitrovic, Izidor Mlakar, Stavros Moiras, Alexander Nussbaumer, Per Öster, Martin Potthast, Marjana Sencar Srdic, Sharikadze Megi, Katerina Slaninová, Benno Stein 0001, Arjen P. de Vries, Vít Vondrák, Saber Zerhoudi
J. Assoc. Inf. Sci. Technol.15
2023 The Infinite Index: Information Retrieval on Generative Text-To-Image Models
abstract
Conditional generative models such as DALL-E and Stable Diffusion generate images based on a user-defined text, the prompt. Finding and refining prompts that produce a desired image has become the art of prompt engineering. Generative models do not provide a built-in retrieval model for a user’s information need expressed through prompts. In light of an extensive literature review, we reframe prompt engineering for generative models as interactive text-based retrieval on a novel kind of “infinite index”. We apply these insights for the first time in a case study on image generation for game design with an expert. Finally, we envision how active learning may help to guide the retrieval of generated images.
Niklas Deckers, Maik Fröbe, Johannes Kiesel, Gianluca Pandolfo, Christopher Schröder 0001, Benno Stein 0001, Martin Potthast
CHIIR7
2023 Overview of PAN 2023: Authorship Verification, Multi-author Writing Style Analysis, Profiling Cryptocurrency Influencers, and Trigger Detection - Extended Abstract
Janek Bevendorff, Mara Chinea-Rios, Marc Franco-Salvador, Annina Heini, Erik Körner, Krzysztof Kredens, Maximilian Mayerl, Piotr Pezik, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Magdalena Wolska, Eva Zangerle
ECIR (3)9
2023 Overview of Touché 2023: Argument and Causal Retrieval - Extended Abstract
Alexander Bondarenko 0001, Maik Fröbe, Johannes Kiesel, Ferdinand Schlatt, Valentin Barrière, Brian Ravenet, Léo Hemamou, Simon Luck, Jan Heinrich Merker, Benno Stein 0001, Martin Potthast, Matthias Hagen
ECIR (3)11
2023 Bootstrapped nDCG Estimation in the Presence of Unjudged Documents
Maik Fröbe, Lukas Gienapp, Martin Potthast, Matthias Hagen
ECIR (1)3
2023 Continuous Integration for Reproducible Shared Tasks with TIRA.io
Maik Fröbe, Matti Wiegmann, Nikolay Kolyada, Bastian Grahm, Theresa Elstner, Frank Loebe, Matthias Hagen, Benno Stein 0001, Martin Potthast
ECIR (3)9
2023 Dynamic Exploratory Search for the Information Retrieval Anthology
Tim Gollub, Jason Brockmeyer, Benno Stein 0001, Martin Potthast
ECIR (3)4
2023 On Stance Detection in Image Retrieval for Argumentation
abstract
Given a text query on a controversial topic, the task of Image Retrieval for Argumentation is to rank images according to how well they can be used to support a discussion on the topic. An important subtask therein is to determine the stance of the retrieved images, i.e., whether an image supports the pro or con side of the topic. In this paper, we conduct a comprehensive reproducibility study of the state of the art as represented by the CLEF'22 Touché lab and an in-house extension of it. Based on the submitted approaches, we developed a unified and modular retrieval process and reimplemented the submitted approaches according to this process. Through this unified reproduction (which also includes models not previously considered), we achieve an effectiveness improvement in argumentative image detection of up to 0.832 [email protected] However, despite this reproduction success, our study also revealed a previously unknown negative result: for stance detection, none of the reproduced or new approaches can convincingly beat a random baseline. To understand the apparent challenges inherent to image stance detection, we conduct a thorough error analysis and provide insight into potential new ways to approach this task.
Miriam Louise Carnot, Lorenz Heinemann, Jan Braker, Tobias Schreieder, Johannes Kiesel, Maik Fröbe, Martin Potthast, Benno Stein 0001
SIGIR7
2023 The Information Retrieval Experiment Platform
abstract
We integrate irdatasets, ir_measures, and PyTerrier with TIRA in the Information Retrieval Experiment Platform (TIREx) to promote more standardized, reproducible, scalable, and even blinded retrieval experiments. Standardization is achieved when a retrieval approach implements PyTerrier's interfaces and the input and output of an experiment are compatible with ir_datasets and ir_measures. However, none of this is a must for reproducibility and scalability, as TIRA can run any dockerized software locally or remotely in a cloud-native execution environment. Version control and caching ensure efficient (re)execution. TIRA allows for blind evaluation when an experiment runs on a remote server or cloud not under the control of the experimenter. The test data and ground truth are then hidden from public access, and the retrieval software has to process them in a sandbox that prevents data leaks.
Maik Fröbe, Jan Heinrich Merker, Sean MacAvaney, Niklas Deckers, Simon Reich, Janek Bevendorff, Benno Stein 0001, Matthias Hagen, Martin Potthast
SIGIR9
2023 The Archive Query Log: Mining Millions of Search Result Pages of Hundreds of Search Engines from 25 Years of Web Archives
abstract
The Archive Query Log (AQL) is a previously unused, comprehensive query log collected at the Internet Archive over the last 25 years. Its first version includes 356 million queries, 137 million search result pages, and 1.4 billion search results across 550 search providers. Although many query logs have been studied in the literature, the search providers that own them generally do not publish their logs to protect user privacy and vital business data. Of the few query logs publicly available, none combines size, scope, and diversity. The AQL is the first to do so, enabling research on new retrieval models and (diachronic) search engine analyses. Provided in a privacy-preserving manner, it promotes open research as well as more transparency and accountability in the search industry.
Jan Heinrich Merker, Sebastian Heineking, Maik Fröbe, Lukas Gienapp, Harrisen Scells, Benno Stein 0001, Matthias Hagen, Martin Potthast
SIGIR8
2023 pybool_ir: A Toolkit for Domain-Specific Search Experiments
abstract
Undertaking research in domain-specific scenarios such as systematic review literature search, legal search, and patent search can often have a high barrier of entry due to complicated indexing procedures and complex Boolean query syntax. Indexing and searching document collections like PubMed in off-the-shelf tools such as Elasticsearch and Lucene often yields less accurate (and less effective) results than the PubMed search engine, i.e., retrieval results do not match what would be retrieved if one issued the same query to PubMed. Furthermore, off-the-shelf tools have their own nuanced query languages and do not allow directly using the often large and complicated Boolean queries seen in domain-specific search scenarios. The pybool_ir toolkit aims to address these problems and to lower the barrier to entry for developing new methods for domain-specific search. The toolkit is an open source package available at https://github.com/hscells/pybool_ir.
Harrisen Scells, Martin Potthast
SIGIR2
2023 Smooth Operators for Effective Systematic Review Queries
abstract
Effective queries are crucial to minimising the time and cost of medical systematic reviews, as all retrieved documents must be judged for relevance. Boolean queries, developed by expert librarians, are the standard for systematic reviews. They guarantee reproducible and verifiable retrieval and more control than free-text queries. However, the result sets of Boolean queries are unranked and difficult to control due to the strict Boolean operators. We address these problems in a single unified retrieval model by formulating a class of smooth operators that are compatible with and extend existing Boolean operators. Our smooth operators overcome several shortcomings of previous extensions of the Boolean retrieval model. In particular, our operators are independent of the underlying ranking function, so that exact-match and large language model rankers can be combined in the same query. We found that replacing Boolean operators with equivalent or similar smooth operators often improves the effectiveness of queries. Their properties make tuning a query to precision or recall intuitive and allow greater control over how documents are retrieved. This additional control leads to more effective queries and reduces the cost of systematic reviews.
Harrisen Scells, Ferdinand Schlatt, Martin Potthast
SIGIR3
2022 Overview of PAN 2022: Authorship Verification, Profiling Irony and Stereotype Spreaders, Style Change Detection, and Trigger Detection - Extended Abstract
Janek Bevendorff, Berta Chulvi, Elisabetta Fersini, Annina Heini, Mike Kestemont, Krzysztof Kredens, Maximilian Mayerl, Reyner Ortega-Bueno, Piotr Pezik, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Magdalena Wolska, Eva Zangerle
ECIR (2)10
2022 Overview of Touché 2022: Argument Retrieval - Extended Abstract
Alexander Bondarenko 0001, Maik Fröbe, Johannes Kiesel, Shahbaz Syed, Timon Ziegenbein, Meriem Beloucif, Alexander Panchenko, Chris Biemann, Benno Stein 0001, Henning Wachsmuth, Martin Potthast, Matthias Hagen
ECIR (2)11
2022 The Power of Anchor Text in the Neural Retrieval Era
Maik Fröbe, Sebastian Günther 0002, Maximilian Probst Gutenberg, Martin Potthast, Matthias Hagen
ECIR (1)4
2022 Visual Web Archive Quality Assessment
Theresa Elstner, Johannes Kiesel, Lars Meyer 0002, Max Martius, Sebastian Heineking, Benno Stein 0001, Martin Potthast
TPDL7
2022 How Train-Test Leakage Affects Zero-Shot Retrieval
Maik Fröbe, Christopher Akiki, Martin Potthast, Matthias Hagen
SPIRE3
2021 Overview of PAN 2021: Authorship Verification, Profiling Hate Speech Spreaders on Twitter, and Style Change Detection - Extended Abstract
Janek Bevendorff, Berta Chulvi, Gretel Liz De la Peña Sarracén, Mike Kestemont, Enrique Manjavacas, Ilia Markov, Maximilian Mayerl, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Magdalena Wolska, Eva Zangerle
ECIR (2)8
2021 Overview of Touché 2021: Argument Retrieval - Extended Abstract
Alexander Bondarenko 0001, Lukas Gienapp, Maik Fröbe, Meriem Beloucif, Yamen Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein 0001, Henning Wachsmuth, Martin Potthast, Matthias Hagen
ECIR (2)10
2021 An Empirical Comparison of Web Page Segmentation Algorithms
Johannes Kiesel, Lars Meyer 0002, Florian Kneist, Benno Stein 0001, Martin Potthast
ECIR (2)5
2021 Identifying Queries in Instant Search Logs
abstract
Query logs of search engines with instant search functionality are challenging for log analysis, since the log entries represent interactions at the keystroke level, rather than at the query level. To enable log analyses at the query level, a user's logged sequence of keystroke-level interactions needs to be mapped to distinct queries. This problem bears strong parallels to session detection in "standard" query logs (i.e., forming groups of subsequent queries on the same topic), but there are salient differences. In this paper, we present a new approach to identifying interactions belonging to the same query in instant query logs. In an experimental comparison, our new approach achieves an F2 score of 0.93 compared to only 0.83 of a state-of-the-art cascading method for query log session detection.
Kristof Komlossy, Benno Stein 0001, Martin Potthast, Matthias Hagen
SIGIR4
2021 CopyCat: Near-Duplicates Within and Between the ClueWeb and the Common Crawl
abstract
The amount of near-duplicates in web crawls like the ClueWeb or Common Crawl demands from their users either to develop a preprocessing pipeline for deduplication, which is costly both computationally and in person hours, or accepting the undesired effects that near-duplicates have on reliability and validity of experiments. We introduce ChatNoir-CopyCat-21, which simplifies deduplication significantly. It comes in two parts: (1) A compilation of near-duplicate documents within the ClueWeb09, the ClueWeb12, and two Common Crawl snapshots, as well as between selections of these crawls, and (2) a software library that implements the deduplication of arbitrary document sets. Our analysis shows that 14--52, of the documents within a crawl and around~0.7--2.5, between the crawls are near-duplicates. Two showcases demonstrate the application and usefulness of our resource.
Maik Fröbe, Janek Bevendorff, Lukas Gienapp, Michael Völske, Benno Stein 0001, Martin Potthast, Matthias Hagen
SIGIR6
2021 The Information Retrieval Anthology
abstract
We present the IR Anthology, a corpus of information retrieval publications accessible via a metadata browser and a full-text search engine. Following the example of the well-known ACL Anthology, the IR Anthology serves as a hub for researchers interested in information retrieval. Our search engine ChatNoir indexes the publications' full texts, enabling a focused search and linking users to the respective publisher's site for personal access. Listing more than 40,000 publications at the time of writing, the IR Anthology can be freely accessed at https://IR.webis.de.
Martin Potthast, Sebastian Günther 0002, Janek Bevendorff, Jan Philipp Bittner, Alexander Bondarenko 0001, Maik Fröbe, Christian Kahmann, Andreas Niekler, Michael Völske, Benno Stein 0001, Matthias Hagen
SIGIR1
2021 Predicting essay quality from search and writing behavior
abstract
Abstract Few studies have investigated how search behavior affects complex writing tasks. We analyze a dataset of 150 long essays whose authors searched the ClueWeb09 corpus for source material, while all querying, clicking, and writing activity was meticulously recorded. We model the effect of search and writing behavior on essay quality using path analysis. Since the boil‐down and build‐up writing strategies identified in previous research have been found to affect search behavior, we model each writing strategy separately. Our analysis shows that the search process contributes significantly to essay quality through both direct and mediated effects, while the author's writing strategy moderates this relationship. Our models explain 25–35% of the variation in essay quality through rather simple search and writing process characteristics alone, a fact that has implications on how search engines could personalize result pages for writing tasks. Authors' writing strategies and associated searching patterns differ, producing differences in essay quality. In a nutshell: essay quality improves if search and writing strategies harmonize—build‐up writers benefit from focused, in‐depth querying, while boil‐down writers fare better with a broader and shallower querying strategy.
Pertti Vakkari, Michael Völske, Martin Potthast, Matthias Hagen, Benno Stein 0001
J. Assoc. Inf. Sci. Technol.3
2021 Meta-Information in Conversational Search
abstract
The exchange of meta-information has always formed part of information behavior. In this article, we show that this rule also extends to conversational search. Information about the user’s information need, their preferences, and the quality of search results are only some of the most salient examples of meta-information that are exchanged as a matter of course in a search conversation. To understand the importance of meta-information for conversational search, we revisit its definition and survey how meta-information has been taken into account in the past in information retrieval. Meta-information has gone by many names, about which a concise overview is provided. An in-depth analysis of the role of meta-information in search and conversation theories reveals that they provide significant support for the importance of meta-information in conversational search. We further identify conversational search datasets are suitable for a deeper inspection with regard to meta-information, namely, Spoken Conversational Search and Microsoft Information-Seeking Conversations. A quantitative data analysis demonstrates the practical significance of meta-information in information-seeking conversations, whereas a qualitative analysis shows the effects of exchanging different types. Finally, we discuss practical applications and challenges of meta-information in conversational search, including a case study of VERSE, an existing search system for the visually impaired.
Johannes Kiesel, Lars Meyer 0002, Martin Potthast, Benno Stein 0001
ACM Trans. Inf. Syst.3
2020 Estimating Topic Difficulty Using Normalized Discounted Cumulated Gain
abstract
Information retrieval evaluation has to consider the varying "difficulty" between topics. Topic difficulty is often defined in terms of the aggregated effectiveness of a set of retrieval systems to satisfy a respective information need. Current approaches to estimate topic difficulty come with drawbacks such as being incomparable across different experimental settings. We introduce a new approach to estimate topic difficulty, which is based on the ratio of systems that achieve an NDCG score that is better than a baseline formed as random ranking of the pool of judged documents. We modify the NDCG measure to explicitly reflect a system's divergence from this hypothetical random ranker. In this way we achieve relative comparability of topic difficulty scores across experimental settings as well as stability to outlier systems?features lacking in previous difficulty estimations. We reevaluate the TREC 2012 Web Track's ad hoc task to demonstrate the feasibility of our approach in practice.
Lukas Gienapp, Benno Stein 0001, Matthias Hagen, Martin Potthast
CIKM4
2020 The Impact of Negative Relevance Judgments on NDCG
abstract
NDCG is one of the most commonly used measures to quantify system performance in retrieval experiments. Though originally not considered, graded relevance judgments nowadays frequently include negative labels. Negative relevance labels cause NDCG to be unbounded. This is probably why widely used implementations of NDCG map negative relevance labels to zero, thus ensuring the resulting scores to originate from the [0,1] range. But zeroing negative labels discards valuable relevance information, e.g., by treating spam documents the same as unjudged ones, which are assigned the relevance label of zero by default. We show that, instead of zeroing negative labels, a min-max-normalization of NDCG retains its statistical power while improving its reliability and stability.
Lukas Gienapp, Maik Fröbe, Matthias Hagen, Martin Potthast
CIKM4
2020 CauseNet: Towards a Causality Graph Extracted from the Web
abstract
Causal knowledge is seen as one of the key ingredients to advance artificial intelligence. Yet, few knowledge bases comprise causal knowledge to date, possibly due to significant efforts required for validation. Notwithstanding this challenge, we compile CauseNet, a large-scale knowledge base of claimed causal relations between causal concepts. By extraction from different semi- and unstructured web sources, we collect more than 11 million causal relations with an estimated extraction precision of 83% and construct the first large-scale and open-domain causality graph. We analyze the graph to gain insights about causal beliefs expressed on the web and we demonstrate its benefits in basic causal question answering. Future work may use the graph for causal reasoning, computational argumentation, multi-hop question answering, and more.
Stefan Heindorf, Yan Scholten, Henning Wachsmuth, Axel-Cyrille Ngonga Ngomo, Martin Potthast
CIKM5
2020 Web Page Segmentation Revisited: Evaluation Framework and Dataset
abstract
Each web page can be segmented into semantically coherent units that fulfill specific purposes. Though the task of automatic web page segmentation was introduced two decades ago, along with several applications in web content analysis, its foundations are still lacking. Specifically, the developed evaluation methods and datasets presume a certain downstream task, which led to a variety of incompatible datasets and evaluation methods. To address this shortcoming, we contribute two resources: (1) An evaluation framework which can be adjusted to downstream tasks by measuring the segmentation similarity regarding visual, structural, and textual elements, and which includes measures for annotator agreement, segmentation quality, and an algorithm for segmentation fusion. (2) The Webis-WebSeg-20 dataset, comprising 42,450~crowdsourced segmentations for 8,490~web pages, outranging existing sources by an order of magnitude. Our results help to better understand the "mental segmentation model'' of human annotators: Among other things we find that annotators mostly agree on segmentations for all kinds of web page elements (visual, structural, and textual). Disagreement exists mostly regarding the right level of granularity, indicating a general agreement on the visual structure of web pages.
Johannes Kiesel, Florian Kneist, Lars Meyer 0002, Kristof Komlossy, Benno Stein 0001, Martin Potthast
CIKM6
2020 Shared Tasks on Authorship Analysis at PAN 2020
Janek Bevendorff, Bilal Ghanem, Anastasia Giahanou, Mike Kestemont, Enrique Manjavacas, Martin Potthast, Francisco M. Rangel Pardo, Paolo Rosso, Günther Specht, Efstathios Stamatatos, Benno Stein 0001, Matti Wiegmann, Eva Zangerle
ECIR (2)6
2020 Touché: First Shared Task on Argument Retrieval
Alexander Bondarenko 0001, Matthias Hagen, Martin Potthast, Henning Wachsmuth, Meriem Beloucif, Chris Biemann, Alexander Panchenko, Benno Stein 0001
ECIR (2)3
2020 The Effect of Content-Equivalent Near-Duplicates on the Evaluation of Search Engines
Maik Fröbe, Jan Philipp Bittner, Martin Potthast, Matthias Hagen
ECIR (2)3
2020 A Search Engine for Police Press Releases to Double-Check the News
Maik Fröbe, Nina Schwanke, Matthias Hagen, Martin Potthast
ECIR (2)4
2020 Sampling Bias Due to Near-Duplicates in Learning to Rank
abstract
Learning to rank~(LTR) is the de facto standard for web search, improving upon classical retrieval models by exploiting (in)direct relevance feedback from user judgments, interaction logs, etc. We investigate for the first time the effect of a sampling bias on LTR~models due to the potential presence of near-duplicate web pages in the training data, and how (in)consistent relevance feedback of duplicates influences an LTR~model's decisions. To examine this bias, we construct a series of specialized LTR~datasets based on the ClueWeb09 corpus with varying amounts of near-duplicates. We devise worst-case and average-case train/test splits that are evaluated on popular pointwise, pairwise, and listwise LTR~models. Our experiments demonstrate that duplication causes overfitting and thus less effective models, making a strong case for the benefits of systematic deduplication before training and model evaluation.
Maik Fröbe, Janek Bevendorff, Jan Heinrich Merker, Martin Potthast, Matthias Hagen
SIGIR4
2020 Abstractive Snippet Generation
abstract
An abstractive snippet is an originally created piece of text to summarize a web page on a search engine results page. Compared to the conventional extractive snippets, which are generated by extracting phrases and sentences verbatim from a web page, abstractive snippets circumvent copyright issues; even more interesting is the fact that they open the door for personalization. Abstractive snippets have been evaluated as equally powerful in terms of user acceptance and expressiveness—but the key question remains: Can abstractive snippets be automatically generated with sufficient quality?
Wei-Fan Chen 0001, Shahbaz Syed, Benno Stein 0001, Matthias Hagen, Martin Potthast
WWW5
2019 Wikipedia Text Reuse: Within and Without
Milad Alshomary, Michael Völske, Tristan Licht, Henning Wachsmuth, Benno Stein 0001, Matthias Hagen, Martin Potthast
ECIR (1)7
2019 A Decade of Shared Tasks in Digital Text Forensics at PAN
Martin Potthast, Paolo Rosso, Efstathios Stamatatos, Benno Stein 0001
ECIR (2)1
2019 Argument Search: Assessing Argument Relevance
abstract
We report on the first user study on assessing argument relevance. Based on a search among more than 300,000 arguments, four standard retrieval models are compared on 40 topics for 20 controversial issues: every issue has one topic with a biased stance and another neutral one. Following TREC, the top results of the different models on a topic were pooled and relevance-judged by one assessor per topic. The assessors also judged the arguments' rhetorical, logical, and dialectical quality, the results of which were cross-referenced with the relevance judgments. Furthermore, the assessors were asked for their personal opinion, and whether it matched the predefined stance of a topic. Among other results, we find that Terrier's implementations of DirichletLM and DPH are on par, significantly outperforming TFIDF and BM25. The judgments of relevance and quality hardly correlate, giving rise to a more diverse set of ranking criteria than relevance alone. We did not measure a significant bias of assessors when their stance is at odds with a topic's stance.
Martin Potthast, Lukas Gienapp, Florian Euchner, Nick Heilenkötter, Nico Weidmann, Henning Wachsmuth, Benno Stein 0001, Matthias Hagen
SIGIR1
2019 Debiasing Vandalism Detection Models at Wikidata
abstract
Crowdsourced knowledge bases like Wikidata suffer from low-quality edits and vandalism, employing machine learning-based approaches to detect both kinds of damage. We reveal that state-of-the-art detection approaches discriminate anonymous and new users: benign edits from these users receive much higher vandalism scores than benign edits from older ones, causing newcomers to abandon the project prematurely. We address this problem for the first time by analyzing and measuring the sources of bias, and by developing a new vandalism detection model that avoids them. Our model FAIR-S reduces the bias ratio of the state-of-the-art vandalism detector WDVD from 310.7 to only 11.9 while maintaining high predictive performance at 0.963 ROC and 0.316 PR.
Stefan Heindorf, Yan Scholten, Gregor Engels, Martin Potthast
WWW4
2019 Modeling the usefulness of search results as measured by information use
Pertti Vakkari, Michael Völske, Martin Potthast, Matthias Hagen, Benno Stein 0001
Inf. Process. Manag.3
2018 Elastic ChatNoir: Search Engine for the ClueWeb and the Common Crawl
Janek Bevendorff, Benno Stein 0001, Matthias Hagen, Martin Potthast
ECIR4
2018 Predicting Retrieval Success Based on Information Use for Writing Tasks
Pertti Vakkari, Michael Völske, Martin Potthast, Matthias Hagen, Benno Stein 0001
TPDL3
2018 A User Study on Snippet Generation: Text Reuse vs. Paraphrases
abstract
The snippets in the result list of a web search engine are built with sentences from the retrieved web pages that match the query. Reusing a web page's text for snippets has been considered fair use under the copyright laws of most jurisdictions. As of recent, notable exceptions from this arrangement include Germany and Spain, where news publishers are entitled to raise claims under a so-called ancillary copyright. A similar legislation is currently discussed at the European Commission. If this development gains momentum, the reuse of text for snippets will soon incur costs, which in turn will give rise to new solutions for generating truly original snippets. A key question in this regard is whether the users will accept any new approach for snippet generation, or whether they will prefer the current model of "reuse snippets." The paper in hand gives a first answer. A crowdsourcing experiment along with a statistical analysis reveals that our test users exert no significant preference for either kind of snippet. Notwithstanding the technological difficulty, this result opens the door to a new snippet synthesis paradigm.
Wei-Fan Chen 0001, Matthias Hagen, Benno Stein 0001, Martin Potthast
SIGIR4
2017 Source Retrieval for Web-Scale Text Reuse Detection
abstract
The first step of text reuse detection addresses the source retrieval problem: given a suspicious document, a set of candidate sources from which text might have been reused have to be retrieved by querying a search engine. Afterwards, in a second step, the retrieved candidates run through a text alignment with the suspicious document in order to identify reused passages. Obviously, any true source of text reuse that is not retrieved during the source retrieval step reduces the overall recall of a reuse detector. Hence, source retrieval is a recall-oriented task, a fact ignored even by experts: Only 3 of 20 teams participating in a respective task at PAN 2012-2016 managed to find more than half of the sources, the best one achieving a recall of only~0.59. We propose a new approach that reaches a recall of~0.89---a performance gain of~51%.
Matthias Hagen, Martin Potthast, Payam Adineh, Ehsan Fatehifar, Benno Stein 0001
CIKM2
2017 Spatio-Temporal Analysis of Reverted Wikipedia Edits
Johannes Kiesel, Martin Potthast, Matthias Hagen, Benno Stein 0001
ICWSM2
2017 A Large-Scale Query Spelling Correction Corpus
abstract
We present a new large-scale collection of 54,772 queries with manually annotated spelling corrections. For 9,170 of the queries (16.74%), spelling variants that are different to the original query are proposed. With its size, our new corpus is an order of magnitude larger than other publicly available query spelling corpora. In addition to releasing the new large-scale corpus, we also provide an implementation of the winner of the Microsoft Speller Challenge from~2011 and compare it on the different publicly available corpora to spelling corrections mined from Google and Bing. This way, we also shed some light on the spelling correction performance of state-of-the-art commercial search systems.
Matthias Hagen, Martin Potthast, Marcel Gohsen, Anja Rathgeber, Benno Stein 0001
SIGIR2
2017 WSDM Cup 2017: Vandalism Detection and Triple Scoring
abstract
The WSDM Cup 2017 was a data mining challenge held in conjunction with the 10th International Conference on Web Search and Data Mining (WSDM). It addressed key challenges of knowledge bases today: quality assurance and entity search. For quality assurance, we tackle the task of vandalism detection, based on a dataset of more than 82 million user-contributed revisions of the Wikidata knowledge base, all of which annotated with regard to whether or not they are vandalism. For entity search, we tackle the task of triple scoring, using a dataset that comprises relevance scores for triples from type-like relations including occupation and country of citizenship, based on about 10,000 human relevance judgments. For reproducibility sake, participants were asked to submit their software on TIRA, a cloud-based evaluation platform, and they were incentivized to share their approaches open source.
Stefan Heindorf, Martin Potthast, Hannah Bast, Björn Buchhold, Elmar Haussmann
WSDM2
2016 How Writers Search: Analyzing the Search and Writing Logs of Non-fictional Essays
abstract
Many writers of non-fictional texts engage intensively in exploratory web search scenarios during their background research on the essay topic. Though understanding such search behavior is necessary for the development of search engines that specifically support writing tasks, it has neither been systematically recorded nor analyzed. This paper contributes part of the missing research: We report on the outcomes of a large-scale corpus construction initiative to acquire detailed interaction logs of writers who were given a writing task on 150 pre-defined TREC topics. The corpus is freely available to foster research on exploratory search. Each essay is at least 5000 words long and comes with a chronological log of search queries, result clicks, web browsing trails, and fine-grained writing revisions that reflect the task completion status. To ensure reproducibility, a fully-fledged, static web search environment has been created on top of the ClueWeb09 corpus as part of our initiative.
Matthias Hagen, Martin Potthast, Michael Völske, Jakob Gomoll, Benno Stein 0001
CHIIR2
2016 Vandalism Detection in Wikidata
abstract
Wikidata is the new, large-scale knowledge base of the Wikimedia Foundation. Its knowledge is increasingly used within Wikipedia itself and various other kinds of information systems, imposing high demands on its integrity. Wikidata can be edited by anyone and, unfortunately, it frequently gets vandalized, exposing all information systems using it to the risk of spreading vandalized and falsified information. In this paper, we present a new machine learning-based approach to detect vandalism in Wikidata. We propose a set of 47 features that exploit both content and context information, and we report on 4 classifiers of increasing effectiveness tailored to this learning task. Our approach is evaluated on the recently published Wikidata Vandalism Corpus WDVC-2015 and it achieves an area under curve value of the receiver operating characteristic, ROC-AUC, of 0.991. It significantly outperforms the state of the art represented by the rule-based Wikidata Abuse Filter (0.865 ROC-AUC) and a prototypical vandalism detector recently introduced by Wikimedia within the Objective Revision Evaluation Service (0.859 ROC-AUC).
Stefan Heindorf, Martin Potthast, Benno Stein 0001, Gregor Engels
CIKM2
2016 Who Wrote the Web? Revisiting Influential Author Identification Research Applicable to Information Retrieval
Martin Potthast, Sarah Braun, Tolga Buz, Fabian Duffhauss, Florian Friedrich, Jörg Marvin Gülzow, Jakob Köhler, Winfried Lötzsch, Maike Elisa Müller, Robert Paßmann, Bernhard Reinke, Lucas Rettenmeier, Thomas Rometsch, Timo Sommer, Michael Träger, Sebastian Wilhelm, Benno Stein 0001, Efstathios Stamatatos, Matthias Hagen
ECIR1
2016 Clickbait Detection
Martin Potthast, Sebastian Köpsel, Benno Stein 0001, Matthias Hagen
ECIR1
2015 Twitter Sentiment Detection via Ensemble Classification Using Averaged Confidence Scores
Matthias Hagen, Martin Potthast, Michel Büchner, Benno Stein 0001
ECIR2
2015 Towards Vandalism Detection in Knowledge Bases: Corpus Construction and Analysis
abstract
We report on the construction of the Wikidata Vandalism Corpus WDVC-2015, the first corpus for vandalism in knowledge bases. Our corpus is based on the entire revision history of Wikidata, the knowledge base underlying Wikipedia. Among Wikidata's 24 million manual revisions, we have identified more than 100,000 cases of vandalism. An in-depth corpus analysis lays the groundwork for research and development on automatic vandalism detection in public knowledge bases. Our analysis shows that 58% of the vandalism revisions can be found in the textual portions of Wikidata, and the remainder in structural content, e.g., subject-predicate-object triples. Moreover, we find that some vandals also target Wikidata content whose manipulation may impact content displayed on Wikipedia, revealing potential vulnerabilities. Given today's importance of knowledge bases for information systems, this shows that public knowledge bases must be used with caution.
Stefan Heindorf, Martin Potthast, Benno Stein 0001, Gregor Engels
SIGIR2
2013 Paraphrase acquisition via crowdsourcing and machine learning
abstract
To paraphrase means to rewrite content while preserving the original meaning. Paraphrasing is important in fields such as text reuse in journalism, anonymizing work, and improving the quality of customer-written reviews. This article contributes to paraphrase acquisition and focuses on two aspects that are not addressed by current research: (1) acquisition via crowdsourcing, and (2) acquisition of passage-level samples. The challenge of the first aspect is automatic quality assurance; without such a means the crowdsourcing paradigm is not effective, and without crowdsourcing the creation of test corpora is unacceptably expensive for realistic order of magnitudes. The second aspect addresses the deficit that most of the previous work in generating and evaluating paraphrases has been conducted using sentence-level paraphrases or shorter; these short-sample analyses are limited in terms of application to plagiarism detection, for example. We present the Webis Crowd Paraphrase Corpus 2011 (Webis-CPC-11), which recently formed part of the PAN 2010 international plagiarism detection competition. This corpus comprises passage-level paraphrases with 4067 positive samples and 3792 negative samples that failed our criteria, using Amazon's Mechanical Turk for crowdsourcing. In this article, we review the lessons learned at PAN 2010, and explain in detail the method used to construct the corpus. The empirical contributions include machine learning experiments to explore if passage-level paraphrases can be identified in a two-class classification problem using paraphrase similarity features, and we find that a k-nearest-neighbor classifier can correctly distinguish between paraphrased and nonparaphrased samples with 0.980 precision at 0.523 recall. This result implies that just under half of our samples must be discarded (remaining 0.477 fraction), but our cost analysis shows that the automation we introduce results in a 18% financial saving and over 100 hours of time returned to the researchers when repeating a similar corpus design. On the other hand, when building an unrelated corpus requiring, say, 25% training data for the automated component, we show that the financial outcome is cost neutral, while still returning over 70 hours of time to the researchers. The work presented here is the first to join the paraphrasing and plagiarism communities.
Steven Burrows, Martin Potthast, Benno Stein 0001
ACM Trans. Intell. Syst. Technol.2
2012 Towards optimum query segmentation: in doubt without
abstract
Query segmentation is the problem of identifying those keywords in a query, which together form compound concepts or phrases like "new york times". Such segments can help a search engine to better interpret a user's intents and to tailor the search results more appropriately. Our contributions to this problem are threefold. (1) We conduct the first large-scale study of human segmentation behavior based on more than 500000 segmentations. (2) We show that the traditionally applied segmentation accuracy measures are not appropriate for such large-scale corpora and introduce new, more robust measures. (3) We develop a new query segmentation approach with the basic idea that, in cases of doubt, it is often better to (partially) leave queries without any segmentation.
Matthias Hagen, Martin Potthast, Anna Beyer, Benno Stein 0001
CIKM2
2012 ChatNoir: a search engine for the ClueWeb09 corpus
abstract
We present the ChatNoir search engine which indexes the entire English part of the ClueWeb09 corpus. Besides Carnegie Mellon's Indri system, ChatNoir is the second publicly available search engine for this corpus. It implements the classic BM25F information retrieval model including PageRank and spam likelihood. The search engine is scalable and returns the first results within three seconds, which is significantly faster than Indri. A convenient API allows for implementing reproducible experiments based on retrieving documents from the ClueWeb09 corpus. The search engine has successfully accomplished a load test involving 100,000 queries.
Martin Potthast, Matthias Hagen, Benno Stein 0001, Jan Graßegger, Maximilian Michel, Martin Tippmann, Clement Welsch
SIGIR1
2012 Information Retrieval in the Commentsphere
abstract
This article studies information retrieval tasks related to Web comments. Prerequisite of such a study and a main contribution of the article is a unifying survey of the research field. We identify the most important retrieval tasks related to comments, namely filtering, ranking, and summarization. Within these tasks, we distinguish two paradigms according to which comments are utilized and which we designate as comment-targeting and comment-exploiting . Within the first paradigm, the comments themselves form the retrieval targets. Within the second paradigm, the commented items form the retrieval targets (i.e., comments are used as an additional information source to improve the retrieval performance for the commented items). We report on four case studies to demonstrate the exploration of the commentsphere under information retrieval aspects: comment filtering, comment ranking, comment summarization and cross-media retrieval. The first three studies deal primarily with comment-targeting retrieval, while the last one deals with comment-exploiting retrieval. Throughout the article, connections to information retrieval research are pointed out.
Martin Potthast, Benno Stein 0001, Fabian Loose, Steffen Becker 0001
ACM Trans. Intell. Syst. Technol.1
2011 Query segmentation revisited
abstract
We address the problem of query segmentation: given a keyword query, the task is to group the keywords into phrases, if possible. Previous approaches to the problem achieve reasonable segmentation performance but are tested only against a small corpus of manually segmented queries. In addition, many of the previous approaches are fairly intricate as they use expensive features and are difficult to be reimplemented.
Matthias Hagen, Martin Potthast, Benno Stein 0001, Christof Bräutigam
WWW2
2010 Cross-Language High Similarity Search: Why No Sub-linear Time Bound Can Be Expected
Maik Anderka, Benno Stein 0001, Martin Potthast
ECIR3
2010 Opinion Summarization of Web Comments
Martin Potthast, Steffen Becker 0001
ECIR1
2010 Netspeak - Assisting Writers in Choosing Words
Martin Potthast, Martin Trenkmann, Benno Stein 0001
ECIR1
2010 Retrieving Customary Web Language to Assist Writers
Benno Stein 0001, Martin Potthast, Martin Trenkmann
ECIR2
2010 The power of naive query segmentation
abstract
We address the problem of query segmentation: given a keyword query submitted to a search engine, the task is to group the keywords into phrases, if possible. Previous approaches to the problem achieve good segmentation performance on a gold standard but are fairly intricate. Our method is easy to implement and comes with a comparable accuracy.
Matthias Hagen, Martin Potthast, Benno Stein 0001, Christof Bräutigam
SIGIR2
2010 Crowdsourcing a wikipedia vandalism corpus
abstract
We report on the construction of the PAN Wikipedia vandalism corpus, PAN-WVC-10, using Amazon's Mechanical Turk. The corpus compiles 32452 edits on 28468 Wikipedia articles, among which 2391 vandalism edits have been identified. 753 human annotators cast a total of 193022 votes on the edits, so that each edit was reviewed by at least 3 annotators, whereas the achieved level of agreement was analyzed in order to label an edit as "regular" or "vandalism." The corpus is available free of charge.
Martin Potthast
SIGIR1
2010 Towards comment-based cross-media retrieval
abstract
This paper investigates whether Web comments can be exploited for cross-media retrieval. Comparing Web items such as texts, images, videos, music, products, or personal profiles can be done at various levels of detail; our focus is on topic similarity. We propose to compare user-supplied comments on Web items in lieu of the commented items themselves. If this approach is feasible, the task of extracting and mapping features between arbitrary pairs of item types can be circumvented, and well-known text retrieval models can be applied instead - given that comments are available. We report on results of a preliminary, but nonetheless large-scale experiment which shows that, if comments on textual items are compared with comments on video items, topically similar pairs achieve a sufficiently high cross-media similarity.
Martin Potthast, Benno Stein 0001, Steffen Becker 0001
WWW1
2009 Measuring the descriptiveness of web comments
abstract
This paper investigates whether Web comments are of descriptive nature, that is, whether the combined text of a set of comments is similar in topic to the commented object. If so, comments may be used in place of the respective object in all kinds of cross-media retrieval tasks. Our experiments reveal that comments on textual objects are indeed descriptive: 10 comments suffice to expect a high similarity between the comments and the commented text; 100-500 comments suffice to replace the commented text in a ranking task, and to measure the contribution of the commenters beyond the commented text.
Martin Potthast
SIGIR1
2008 A Wikipedia-Based Multilingual Retrieval Model
Martin Potthast, Benno Stein 0001, Maik Anderka
ECIR1
2008 Automatic Vandalism Detection in Wikipedia
Martin Potthast, Benno Stein 0001, Robert Gerling
ECIR1
2007 Wikipedia in the pocket: indexing technology for near-duplicate detection and high similarity search
abstract
We develop and implement a new indexing technology which allows us to use complete (and possibly very large) documents as queries, while having a retrieval performance comparable to a standard term query. Our approach aims at retrieval tasks such as near duplicate detection and high similarity search. To demonstrate the performance of our technology we have compiled the search index "Wikipedia in the Pocket", which contains about 2 million English and German Wikipedia articles.1 This index--along with a search interface--fits on a conventional CD (0.7 gigabyte). The ingredients of our indexing technology are similarity hashing and minimal perfect hashing.
Martin Potthast
SIGIR1
2007 Strategies for retrieving plagiarized documents
abstract
For the identification of plagiarized passages in large document collections we present retrieval strategies which rely on stochastic sampling and chunk indexes. Using the entire Wikipedia corpus we compile n-gram indexes and compare them to a new kind of fingerprint index in a plagiarism analysis use case. Our index provides an analysis speed-up by factor 1.5 and is an order of magnitude smaller, while being equivalent in terms of precision and recall.
Benno Stein 0001, Sven Meyer zu Eissen, Martin Potthast
SIGIR3