Alexander Bondarenko 0001

dblp:234/7521 · DBLP profile ↗
← Back
15ranked-venue papers in the field
9as first author
12since 2021 · last 2025
0000-0002-1678-0094ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 13 (7 first)Data Mining & Knowledge Discovery · 2 (2 first)
YearPublicationVenuePosition
2025 Axiomatic Re-Ranking for Argument Retrieval
abstract
Information retrieval axioms are formalized constraints that retrieval systems should ideally satisfy (e.g., to rank documents higher that contain the query terms more often). In this paper, we propose new axioms that focus on the scenario of argument retrieval: retrieval for queries that need arguments in the results. Our underlying axiomatic idea is that in such scenarios, documents should be prioritized with argumentative units that are similar to the query. We test our new axioms in re-ranking experiments on the data of the Touché ~2020 and~2021 shared task on argument retrieval for controversial questions, and show that the new axioms can improve the effectiveness of Touché's strong DirichletLM baseline model and even of the top-performing system from Touché ~2021, a system already specifically optimized for argument retrieval. Finally, we also propose a new method for visualizing the relationships between axioms based on their effects in re-ranking settings.
Maximilian Heinrich, Marvin Vogel, Alexander Bondarenko 0001, Matthias Hagen, Benno Stein 0001
SIGIR3
2024 Systematic Evaluation of Neural Retrieval Models on the Touché 2020 Argument Retrieval Subset of BEIR
abstract
The zero-shot effectiveness of neural retrieval models is often evaluated on the BEIR benchmark---a combination of different IR evaluation datasets. Interestingly, previous studies found that particularly on the BEIR~subset Touché 2020, an argument retrieval task, neural retrieval models are considerably less effective than BM25. Still, so far, no further investigation has been conducted on what makes argument retrieval so "special''. To more deeply analyze the respective potential limits of neural retrieval models, we run a reproducibility study on the Touché 2020 data. In our study, we focus on two experiments: (i) a black-box evaluation (i.e., no model retraining), incorporating a theoretical exploration using retrieval axioms, and (ii) a data denoising evaluation involving post-hoc relevance judgments. Our black-box evaluation reveals an inherent bias of neural models towards retrieving short passages from the Touché 2020 data, and we also find that quite a few of the neural models' results are unjudged in the Touché 2020 data. As many of the short Touché passages are not argumentative and thus non-relevant per se, and as the missing judgments complicate fair comparison, we denoise the Touché 2020 data by excluding very short passages (less than 20 words) and by augmenting the unjudged data with post-hoc judgments following the Touché guidelines. On the denoised data, the effectiveness of the neural models improves by up to 0.52 in nDCG@10, but BM25 is still more effective. Our code and the augmented Touché 2020 dataset are available at https://github.com/castorini/touche-error-analysis.
Nandan Thakur, Luiz Bonifacio, Maik Fröbe, Alexander Bondarenko 0001, Ehsan Kamalloo, Martin Potthast, Matthias Hagen, Jimmy Lin
SIGIR4
2023 Overview of Touché 2023: Argument and Causal Retrieval - Extended Abstract
Alexander Bondarenko 0001, Maik Fröbe, Johannes Kiesel, Ferdinand Schlatt, Valentin Barrière, Brian Ravenet, Léo Hemamou, Simon Luck, Jan Heinrich Merker, Benno Stein 0001, Martin Potthast, Matthias Hagen
ECIR (3)1
2023 Consumer Health Question Answering Using Off-the-Shelf Components
Alexander Pugachev, Ekaterina Artemova, Alexander Bondarenko 0001, Pavel Braslavski 0001
ECIR (2)3
2022 A User Study on Clarifying Comparative Questions
abstract
Vague or ambiguous queries can make it difficult for a search engine to correctly interpret a user’s underlying information need. A relatively “simple” solution then is result diversification to cover different interpretations, while in more “conversational” search interfaces, the user can be prompted to clarify their original request. We study clarification in the scenario of comparative questions that ask to compare several options. In our experiment that reflects a conversational search interface with a clarification component, 70% of the study participants find clarifications useful to retrieve relevant results for questions with unclear comparison aspects (e.g., “Which is better, Bali or Phuket?”) or without explicit comparison objects and aspects (e.g., “What is the best antibiotic?”).
Alexander Bondarenko 0001, Ekaterina Shirshakova, Matthias Hagen
CHIIR1
2022 Overview of Touché 2022: Argument Retrieval - Extended Abstract
Alexander Bondarenko 0001, Maik Fröbe, Johannes Kiesel, Shahbaz Syed, Timon Ziegenbein, Meriem Beloucif, Alexander Panchenko, Chris Biemann, Benno Stein 0001, Henning Wachsmuth, Martin Potthast, Matthias Hagen
ECIR (2)1
2022 Identifying Argumentative Questions in Web Search Logs
abstract
We present an approach to identify argumentative questions among web search queries. Argumentative questions ask for reasons to support a certain stance on a controversial topic, such as ''Should marijuana be legalized?'' Controversial topics entail opposing stances, and hence can be supported or opposed by various arguments. Argumentative questions pose a challenge for search engines since they should be answered with both pro and con arguments in order to not bias a user toward a certain stance.
Yamen Ajjour, Pavel Braslavski 0001, Alexander Bondarenko 0001, Benno Stein 0001
SIGIR3
2022 Axiomatic Retrieval Experimentation with ir_axioms
abstract
Axiomatic approaches to information retrieval have played a key role in determining basic constraints that characterize good retrieval models. Beyond their importance in retrieval theory, axioms have been operationalized to improve an initial ranking, to "guide" retrieval, or to explain some model's rankings. However, recent open-source retrieval frameworks like PyTerrier and Pyserini, which made it easy to experiment with sparse and dense retrieval models, have not included any retrieval axiom support so far.
Alexander Bondarenko 0001, Maik Fröbe, Jan Heinrich Merker, Benno Stein 0001, Michael Völske, Matthias Hagen
SIGIR1
2022 Towards Understanding and Answering Comparative Questions
abstract
In this paper, we analyze comparative questions and answers. At least 3%~of the questions submitted to search engines are comparative; ranging from simple facts like "Did Messi or Ronaldo score more goals in 2021?'' to life-changing and probably highly subjective questions like "Is it better to move abroad or stay?''. Ideally, answers to subjective comparative questions would reflect diverse opinions so that the asker can come to a well-informed decision. To better understand the information needs behind comparative questions, we develop approaches to extract the mentioned comparison objects and aspects. As a first step to answer comparative questions, we develop an approach that detects the stances of potential result nuggets (i.e., text passages containing the comparison objects). Our approaches are trained and evaluated on a set of 31,000~English questions from existing datasets that we label as comparative or not. In the 3,500~comparative questions, we label the comparison objects, aspects, and predicates. For 950~questions, we collect answers from online forums and label the stance towards the comparison objects. In the experiments, our approaches recall~71% of the comparative questions with a perfect precision of~1.0, recall~92% of subjective comparative questions with a precision of~0.98, and identify the comparison objects and aspects with an F1 of~0.93 and~0.80, respectively. The stance detector fine-tuned on pairs of objects and answers achieves an accuracy of~0.63.
Alexander Bondarenko 0001, Yamen Ajjour, Valentin Dittmar, Niklas Homann, Pavel Braslavski 0001, Matthias Hagen
WSDM1
2021 Misbeliefs and Biases in Health-Related Searches
abstract
Quality of search engine results returned to health-related questions is very critical, since a searcher may directly trust any suggestion in the top results. We analyze search questions that mention diseases / symptoms and remedies that are potential health-related misbeliefs. Using lists of medical and alternative medicine terms, we extract health-related search questions from 1.5~billion questions submitted to Yandex. As an initial study, we sample 30 frequent questions that contain a disease--remedy pair like "Can hepatitis be cured with milk thistle?". For each question, we carefully identify a ground truth answer in the medical literature and annotate the top-10 Yandex search result snippets as confirming the belief, rejecting it, or giving no answer. Our analysis shows that about 44%~of the snippets (that users may simply interpret as definitive answers!) confirm some untrue beliefs and are wrong, and only few include health risk warnings about using toxic plants.
Alexander Bondarenko 0001, Ekaterina Shirshakova, Marina Driker, Matthias Hagen, Pavel Braslavski 0001
CIKM1
2021 Overview of Touché 2021: Argument Retrieval - Extended Abstract
Alexander Bondarenko 0001, Lukas Gienapp, Maik Fröbe, Meriem Beloucif, Yamen Ajjour, Alexander Panchenko, Chris Biemann, Benno Stein 0001, Henning Wachsmuth, Martin Potthast, Matthias Hagen
ECIR (2)1
2021 The Information Retrieval Anthology
abstract
We present the IR Anthology, a corpus of information retrieval publications accessible via a metadata browser and a full-text search engine. Following the example of the well-known ACL Anthology, the IR Anthology serves as a hub for researchers interested in information retrieval. Our search engine ChatNoir indexes the publications' full texts, enabling a focused search and linking users to the respective publisher's site for personal access. Listing more than 40,000 publications at the time of writing, the IR Anthology can be freely accessed at https://IR.webis.de.
Martin Potthast, Sebastian Günther 0002, Janek Bevendorff, Jan Philipp Bittner, Alexander Bondarenko 0001, Maik Fröbe, Christian Kahmann, Andreas Niekler, Michael Völske, Benno Stein 0001, Matthias Hagen
SIGIR5
2020 Touché: First Shared Task on Argument Retrieval
Alexander Bondarenko 0001, Matthias Hagen, Martin Potthast, Henning Wachsmuth, Meriem Beloucif, Chris Biemann, Alexander Panchenko, Benno Stein 0001
ECIR (2)1
2020 Comparative Web Search Questions
abstract
\beginabstract We analyze comparative questions, i.e., questions asking to compare different items, that were submitted to Yandex in 2012. Responses to such questions might be quite different from the simple "ten blue links'' and could, for example, aggregate pros and cons of the different options as direct answers. However, changing the result presentation is an intricate decision such that the classification of comparative questions forms a highly precision-oriented task.
Alexander Bondarenko 0001, Pavel Braslavski 0001, Michael Völske, Rami Aly, Maik Fröbe, Alexander Panchenko, Chris Biemann, Benno Stein 0001, Matthias Hagen
WSDM1
2019 Answering Comparative Questions: Better than Ten-Blue-Links?
abstract
We present CAM (comparative argumentative machine), a novel open-domain IR system to argumentatively compare objects with respect to information extracted from the Common Crawl. In a user study, the participants obtained 15% more accurate answers using CAM compared to a "traditional" keyword-based search and were 20% faster in finding the answer to comparative questions.
Matthias Schildwächter, Alexander Bondarenko 0001, Julian Zenker, Matthias Hagen, Chris Biemann, Alexander Panchenko
CHIIR2