VLDB 2026 Research / reviewers in the wild / expert
Johannes Kiesel
dblp:118/3606
· DBLP profile ↗
27ranked-venue papers in the field
12as first author
22since 2021 · last 2026
0000-0002-1617-6508ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 27 (12 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Does Cognitive Load Affect Human Accuracy in Detecting Voice-Based Deepfakes?abstractDeepfake technologies are powerful tools that can be misused for malicious purposes such as spreading disinformation on social media. The effectiveness of such malicious applications depends on the ability of deepfakes to deceive their audience. Therefore, researchers have investigated human abilities to detect deepfakes in various studies. However, most of these studies were conducted with participants who focused exclusively on the detection task; hence the studies may not provide a complete picture of human abilities to detect deepfakes under realistic conditions: Social media users are exposed to cognitive load on the platform, which can impair their detection abilities. In this paper, we investigate the influence of cognitive load on human detection abilities of voice-based deepfakes in an empirical study with 30 participants. Our results suggest that low cognitive load does not generally impair detection abilities, and that the simultaneous exposure to a secondary stimulus can actually benefit people in the detection task. Marcel Gohsen, Nicola Lea Libera, Johannes Kiesel, Jan Ehlers 0001, Benno Stein 0001 |
CHIIR | 3 |
| 2026 | Overview of Touché 2026: Argumentation Systems - Extended Abstract
Johannes Kiesel, Marc Feger, Tim Hagen, Sebastian Heineking, Maximilian Heinrich, Maik Fröbe, Katarina Boland, Wilhelm Pertsch, Julia Romberg, Ines Zelch, Stefan Dietze, Matthias Hagen, Martin Potthast, Benno Stein 0001 |
ECIR (4) | 1 |
| 2026 | TREC iKAT 2025: A Test Collection for the Offline and Interactive Evaluation of Conversational SearchabstractConversational search agents, especially with the advent of large language models, have developed into useful tools to satisfy complex information needs of their users. Former research has shown that personalization (i.e., adaptation of agent responses to the preferences and traits of the user) can increase the relevance and perceived answer quality of these systems even further. However, developing accurate personalization methods typically requires rich datasets, both in terms of user profiles and complex conversations, for which only a few resources are publicly available. Over the past three years, the goal of the TREC Interactive Knowledge Assistance Track (iKAT) has been to bridge this gap. In organizing this shared task, we have developed a collection of complex information needs and associated conversations, made to challenge today's conversational agents and thus highlight aspects in need of further research. In this paper, we present the resources made for iKAT 2025, focusing on multi-session conversations (i.e., multiple dialogues per user), dynamically evolving user models, mixed-initiative dialogues, and large-scale human and automatic assessments. In addition to manually designed user profiles and conversations, the test collection for 2025 also contains dialogues between participating systems and our user simulators. All the resources are publicly available in our repository, including the system evaluation both as a result of the offline (i.e., test collection-based) and interactive tasks (i.e., user simulation-based), as well as their source code and model weights, to foster future research in this direction. Zahra Abbasiantaeb, Simon Lupart, Marcel Gohsen, Nailia Mirzakhmedova, Johannes Kiesel, Jeff Dalton 0001, Mohammad Aliannejadi |
SIGIR | 5 |
| 2026 | Sim.API: A Middleware to Simplify the Use of User Simulators for Shared Tasks in Conversational Search
Marcel Gohsen, Nailia Mirzakhmedova, Zahra Abbasiantaeb, Johannes Kiesel, Simon Lupart, Jeff Dalton 0001, Benno Stein 0001, Mohammad Aliannejadi |
SIGIR | 4 |
| 2026 | Do Simulated Users Need to Remember? Analyzing the Impact of Memory Models in Conversational Search EvaluationabstractConversational search systems are typically evaluated using a fixed reference collection of conversations or through user studies with a live system. However, fixed-reference conversations can cover only a few plausible conversations, and user studies are costly, time-consuming, and often hard to reproduce. A promising alternative that avoids coverage and cost issues is user simulation, in which a computer program takes on the role of a user and interacts with the system under evaluation. But the complexity of human search behavior raises the question of how ''realistic'' the simulations actually need to be for reliable evaluations of conversational search systems. In this paper, we ask: Do simulated users need to remember? While real users may learn and forget information during conversational search sessions, which inspired previous research to also model memory capabilities in simulations, it remains unclear whether this actually influences the results of system evaluations. To investigate the impact of memory modeling, we analyze conversations of simulated users and of humans with four conversational search systems. Our results suggest that incorporating long-term memory into simulators can help reproduce system effectiveness rankings obtained from human conversations, whereas incorporating short-term memory can diminish the reproduction. We also find that simulators are generally valid and reproducible---and memory modeling even increases run-to-run reproducibility of system rankings---but overall, simulations approximate human evaluation scores better when ''helpful'' assistants are evaluated than when assistants with deteriorated response quality are assessed. Our code and data are available at https://github.com/webis-de/SIGIR-26. Nailia Mirzakhmedova, Marcel Gohsen, Johannes Kiesel, Matthias Hagen, Benno Stein 0001 |
SIGIR | 3 |
| 2025 | ImageCLEF 2025: Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications
Bogdan Ionescu, Henning Müller, Dan-Cristian Stanciu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Andrea M. Storås, Asma Ben Abacha, Benjamin Bracke, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Diandra Fabre, Didier Schwab, Dimitar Dimitrov 0003, Emmanuelle Esperança-Rodier, Mihai Gabriel Constantin, Helmut Becker, Hendrik Damm, Henning Schäfer, Ivan Rodkin, Ivan Koychev, Johannes Kiesel, Johannes Rückert, Josep Malvehy, Liviu-Daniel Stefan, Louise Bloch, Martin Potthast, Maximilian Heinrich, Michael Riegler 0001, Mihai Dogariu, Noel Codella, Pål Halvorsen, Preslav Nakov, Raphael Brüngel, Roberto A. Novoa, Rocktim Jyoti Das, Steven Alexander Hicks, Sushant Gautam, Tabea Margareta Grace Pakull, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim, Zhuohan Xie |
ECIR (5) | 26 |
| 2025 | Overview of Touché 2025: Argumentation Systems - Extended Abstract
Johannes Kiesel, Çagri Çöltekin, Marcel Gohsen, Sebastian Heineking, Maximilian Heinrich, Maik Fröbe, Tim Hagen, Mohammad Aliannejadi, Tomaz Erjavec, Matthias Hagen, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Harrisen Scells, Ines Zelch, Martin Potthast, Benno Stein 0001 |
ECIR (5) | 1 |
| 2024 | The Eighth Workshop on Search-Oriented Conversational Artificial Intelligence (SCAI'24)abstractWith the emergence of voice assistants and large language models, conversational interaction with information has become part of everyday life. The eighth edition of the search-oriented conversational AI (SCAI) workshop brings together practitioners and researchers from various disciplines to discuss challenges and advances in conversational search systems. This year’s edition focuses on evaluations beyond relevance and accuracy and looks at conversational search from the user’s perspective. The workshop features a shared task on user-centered evaluation datasets and metrics, challenging participants to develop new and innovative ways to evaluate conversational search systems while accounting for the needs and preferences of users. Alexander Frummet, Andrea Papenmeier, Maik Fröbe, Johannes Kiesel |
CHIIR | 4 |
| 2024 | Advancing Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications with ImageCLEF 2024
Bogdan Ionescu, Henning Müller, Ana-Maria Claudia Dragulinescu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandru Stan, Andrea M. Storås, Asma Ben Abacha, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Didier Schwab, Emmanuelle Esperança-Rodier, George Ioannidis, Griffin Adams, Henning Schäfer, Hugo Manguinhas, Ioan Coman, Johanna Schöler, Johannes Kiesel, Johannes Rückert, Louise Bloch, Martin Potthast, Maximilian Heinrich, Meliha Yetisgen, Michael Riegler 0001, Neal Snider, Pål Halvorsen, Raphael Brüngel, Steven Alexander Hicks, Vajira Thambawita, Vassili Kovalev, Yuri Prokopchuk, Wen-Wai Yim |
ECIR (6) | 24 |
| 2024 | Overview of Touché 2024: Argumentation Systems
Johannes Kiesel, Çagri Çöltekin, Maximilian Heinrich, Maik Fröbe, Milad Alshomary, Bertrand De Longueville, Tomaz Erjavec, Nicolas Handke, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Theresa Reitis-Münstermann, Mario Scharfbillig, Nicolas Stefanovitch, Henning Wachsmuth, Martin Potthast, Benno Stein 0001 |
ECIR (5) | 1 |
| 2024 | Simulating Follow-Up Questions in Conversational Search
Johannes Kiesel, Marcel Gohsen, Nailia Mirzakhmedova, Matthias Hagen, Benno Stein 0001 |
ECIR (2) | 1 |
| 2024 | Evaluating Generative Ad Hoc Information RetrievalabstractRecent advances in large language models have enabled the development of viable generative retrieval systems. Instead of a traditional document ranking, generative retrieval systems often directly return a grounded generated text as a response to a query. Quantifying the utility of the textual responses is essential for appropriately evaluating such generative ad hoc retrieval. Yet, the established evaluation methodology for ranking-based ad hoc retrieval is not suited for the reliable and reproducible evaluation of generated responses. To lay a foundation for developing new evaluation methods for generative retrieval systems, we survey the relevant literature from the fields of information retrieval and natural language processing, identify search tasks and system architectures in generative retrieval, develop a new user model, and study its operationalization. Lukas Gienapp, Harrisen Scells, Niklas Deckers, Janek Bevendorff, Shuai Wang 0032, Johannes Kiesel, Shahbaz Syed, Maik Fröbe, Guido Zuccon, Benno Stein 0001, Matthias Hagen, Martin Potthast |
SIGIR | 6 |
| 2023 | The Infinite Index: Information Retrieval on Generative Text-To-Image ModelsabstractConditional generative models such as DALL-E and Stable Diffusion generate images based on a user-defined text, the prompt. Finding and refining prompts that produce a desired image has become the art of prompt engineering. Generative models do not provide a built-in retrieval model for a user’s information need expressed through prompts. In light of an extensive literature review, we reframe prompt engineering for generative models as interactive text-based retrieval on a novel kind of “infinite index”. We apply these insights for the first time in a case study on image generation for game design with an expert. Finally, we envision how active learning may help to guide the retrieval of generated images. Niklas Deckers, Maik Fröbe, Johannes Kiesel, Gianluca Pandolfo, Christopher Schröder 0001, Benno Stein 0001, Martin Potthast |
CHIIR | 3 |
| 2023 | Guiding Oral Conversations: How to Nudge Users Towards Asking Questions?abstractHow could an envisioned voice-based conversational information system assist the information seeker when the seeker does not know how to continue the conversation? The system could explicitly suggest a question to ask after each of its responses, but this approach quickly feels restrictive, repetitive, and interrupts immersion in the conversation. In this paper, we explore, for the first time, unobtrusive syntactic and auditive modifications of oral system responses to nudge information seekers towards asking about specific topics. We report the results of a crowdsourcing study with 965 participations that investigated the effectiveness and drawbacks of different modifications in three information scenarios. Marcel Gohsen, Johannes Kiesel, Mariam Korashi, Jan Ehlers 0001, Benno Stein 0001 |
CHIIR | 2 |
| 2023 | Overview of Touché 2023: Argument and Causal Retrieval - Extended Abstract
Alexander Bondarenko 0001, Maik Fröbe, Johannes Kiesel, Ferdinand Schlatt, Valentin Barrière, Brian Ravenet, Léo Hemamou, Simon Luck, Jan Heinrich Merker, Benno Stein 0001, Martin Potthast, Matthias Hagen |
ECIR (3) | 3 |
| 2023 | An Empirical Comparison of Web Content Extraction AlgorithmsabstractMain content extraction from web pages-sometimes also called boilerplate removal-has been a research topic for over two decades. Yet despite web pages being delivered in a machine-readable markup format, extracting the actual content is still a challenge today. Even with the latest HTML5 standard, which defines many semantic elements to mark content areas, web page authors do not always use semantic markup correctly or to its full potential, making it hard for automated systems to extract the relevant information. A high-precision, high-recall content extraction is crucial for downstream applications such as search engines, AI language tools, distraction-free reader modes in users' browsers, and other general assistive technologies. For such a fundamental task, however, surprisingly few openly available extraction systems or training and benchmarking datasets exist. Even less research has gone into the rigorous evaluation and a true apples-to-apples comparison of the few extraction systems that do exist. To get a better grasp on the current state of the art in the field, we combine and clean eight existing human-labeled web content extraction datasets. On the combined dataset, we evaluate 14~competitive main content extraction systems and five baseline approaches. Finally, we build three ensembles as new state-of-the-art extraction baselines. We find that the performance of existing systems is quite genre-dependent and no single extractor performs best on all types of web pages. Janek Bevendorff, Sanket Gupta, Johannes Kiesel, Benno Stein 0001 |
SIGIR | 3 |
| 2023 | On Stance Detection in Image Retrieval for ArgumentationabstractGiven a text query on a controversial topic, the task of Image Retrieval for Argumentation is to rank images according to how well they can be used to support a discussion on the topic. An important subtask therein is to determine the stance of the retrieved images, i.e., whether an image supports the pro or con side of the topic. In this paper, we conduct a comprehensive reproducibility study of the state of the art as represented by the CLEF'22 Touché lab and an in-house extension of it. Based on the submitted approaches, we developed a unified and modular retrieval process and reimplemented the submitted approaches according to this process. Through this unified reproduction (which also includes models not previously considered), we achieve an effectiveness improvement in argumentative image detection of up to 0.832 [email protected] However, despite this reproduction success, our study also revealed a previously unknown negative result: for stance detection, none of the reproduced or new approaches can convincingly beat a random baseline. To understand the apparent challenges inherent to image stance detection, we conduct a thorough error analysis and provide insight into potential new ways to approach this task. Miriam Louise Carnot, Lorenz Heinemann, Jan Braker, Tobias Schreieder, Johannes Kiesel, Maik Fröbe, Martin Potthast, Benno Stein 0001 |
SIGIR | 5 |
| 2022 | What is That? Crowdsourcing Questions to a Virtual ExhibitionabstractVirtual environments with an ambient natural interface that allows to retrieve information for learning about the environment are a promising combination for implementing engaging virtual exhibitions. As a step towards better understanding search behavior in such exhibitions, this paper contributes the data from an exploratory study with participants asking questions on a real-world historical room, the Gropiuszimmer at the Weimar Bauhaus, while being on an “online virtual tour” through the room. The dataset comprises 849 manually categorized questions (557 in English, 292 in German) from 63 participants combined with a detailed interaction log, which allows replaying each session (29 hours total). The presented dataset and analyses aim to provide researchers and practitioners with a starting point to develop in-depth studies and prototypical systems. Johannes Kiesel, Volker Bernhard, Marcel Gohsen, Josef Roth, Benno Stein 0001 |
CHIIR | 1 |
| 2022 | Overview of Touché 2022: Argument Retrieval - Extended Abstract
Alexander Bondarenko 0001, Maik Fröbe, Johannes Kiesel, Shahbaz Syed, Timon Ziegenbein, Meriem Beloucif, Alexander Panchenko, Chris Biemann, Benno Stein 0001, Henning Wachsmuth, Martin Potthast, Matthias Hagen |
ECIR (2) | 3 |
| 2022 | Visual Web Archive Quality Assessment
Theresa Elstner, Johannes Kiesel, Lars Meyer 0002, Max Martius, Sebastian Heineking, Benno Stein 0001, Martin Potthast |
TPDL | 2 |
| 2021 | An Empirical Comparison of Web Page Segmentation Algorithms
Johannes Kiesel, Lars Meyer 0002, Florian Kneist, Benno Stein 0001, Martin Potthast |
ECIR (2) | 1 |
| 2021 | Meta-Information in Conversational SearchabstractThe exchange of meta-information has always formed part of information behavior. In this article, we show that this rule also extends to conversational search. Information about the user’s information need, their preferences, and the quality of search results are only some of the most salient examples of meta-information that are exchanged as a matter of course in a search conversation. To understand the importance of meta-information for conversational search, we revisit its definition and survey how meta-information has been taken into account in the past in information retrieval. Meta-information has gone by many names, about which a concise overview is provided. An in-depth analysis of the role of meta-information in search and conversation theories reveals that they provide significant support for the importance of meta-information in conversational search. We further identify conversational search datasets are suitable for a deeper inspection with regard to meta-information, namely, Spoken Conversational Search and Microsoft Information-Seeking Conversations. A quantitative data analysis demonstrates the practical significance of meta-information in information-seeking conversations, whereas a qualitative analysis shows the effects of exchanging different types. Finally, we discuss practical applications and challenges of meta-information in conversational search, including a case study of VERSE, an existing search system for the visually impaired. Johannes Kiesel, Lars Meyer 0002, Martin Potthast, Benno Stein 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2020 | Investigating Expectations for Voice-based and Conversational Argument Search on the WebabstractMillions of arguments are shared on the web. Future information systems will be able to exploit this valuable knowledge source and to retrieve arguments relevant and convincing to our specific need---all with an interface as intuitive as asking your friend "Why ...". Although recent advancements in argument mining, conversational search, and voice recognition have put such systems within reach, many questions remain open, especially on the interface side. In this regard the paper at hand presents the first study of argument search behavior. We conduct an online-survey and a focused user study, putting emphasis on what people expect argument search to be like, rather than on what current first-generation systems provide. Our participants expected to use voice-based argument search mostly at home, but also together with others. Moreover, they expect such search systems to provide rich information on retrieved arguments, such as the source, supporting evidence, and background knowledge on entities or events mentioned. In observed interactions with a simulated system we found that the participants adapted their search behavior to different types of tasks, and that up-front categorization of the retrieved arguments is perceived as helpful if this is short. Our findings are directly applicable to the design of argument search systems, not only voice-based ones. Johannes Kiesel, Kevin Lang, Henning Wachsmuth, Eva Hornecker, Benno Stein 0001 |
CHIIR | 1 |
| 2020 | Web Page Segmentation Revisited: Evaluation Framework and DatasetabstractEach web page can be segmented into semantically coherent units that fulfill specific purposes. Though the task of automatic web page segmentation was introduced two decades ago, along with several applications in web content analysis, its foundations are still lacking. Specifically, the developed evaluation methods and datasets presume a certain downstream task, which led to a variety of incompatible datasets and evaluation methods. To address this shortcoming, we contribute two resources: (1) An evaluation framework which can be adjusted to downstream tasks by measuring the segmentation similarity regarding visual, structural, and textual elements, and which includes measures for annotator agreement, segmentation quality, and an algorithm for segmentation fusion. (2) The Webis-WebSeg-20 dataset, comprising 42,450~crowdsourced segmentations for 8,490~web pages, outranging existing sources by an order of magnitude. Our results help to better understand the "mental segmentation model'' of human annotators: Among other things we find that annotators mostly agree on segmentations for all kinds of web page elements (visual, structural, and textual). Disagreement exists mostly regarding the right level of granularity, indicating a general agreement on the visual structure of web pages. Johannes Kiesel, Florian Kneist, Lars Meyer 0002, Kristof Komlossy, Benno Stein 0001, Martin Potthast |
CIKM | 1 |
| 2019 | Clarifying False Memories in Voice-based SearchabstractQueries containing false memories (i.e., attributes the user misremembered about a searched item) represent a challenge for search systems. A query with a false memory will match inadequate results or even no result, and an automatic query correction is necessary to satisfy the user expectations. For voice-based search interfaces, which aim at a natural, dialog-based search experience, a sensible answer to this kind of unintentionally ill-posed queries is even more crucial. However, the usual solutions in display-based interfaces for queries without matches (e.g., suggesting to drop some query terms) cannot really be transferred to the voice-based setting. Based on the assumption that false memory queries could be identified---a research problem in its own right---, we present the first user study on how voice-based search systems may communicate the respective corrections to a user. Our study compares the user satisfaction in a voice-based search setting for three kinds of false memory clarifications and a baseline case where the system just answers "I don't know.'' Our findings suggest that (1)~users are more satisfied when they receive a clarification that and how the system corrected a false memory, (2)~users even prefer failed correction attempts over no such attempt, and (3)~the tone of the clarification has to be considered for the best possible user satisfaction as well. Johannes Kiesel, Arefeh Bahrami, Benno Stein 0001, Avishek Anand, Matthias Hagen |
CHIIR | 1 |
| 2018 | Toward Voice Query ClarificationabstractQuery suggestions are a standard means to clarify the intent of underspecified queries. In a voice-based search setting, the compilation of query suggestions is not straightforward, and user-centric research targeting query underspecification is lacking so far. Our paper analyses a specific type of ambiguous voice queries and studies the impact of various kinds of voice query clarifications offered by the system and its impact on user satisfaction. We conduct a user study that measures the satisfaction for clarifications that are explicitly invoked and presented by seven different methods. Our findings include that (1) user experience depends on language proficiency levels, (2) users are not dissatisfied when prompted for clarifications (in fact, enjoy it sometimes), and (3) the most effective way of query clarification depends on the number and lengths of the possible answers. Johannes Kiesel, Arefeh Bahrami, Benno Stein 0001, Avishek Anand, Matthias Hagen |
SIGIR | 1 |
| 2017 | Spatio-Temporal Analysis of Reverted Wikipedia Edits
Johannes Kiesel, Martin Potthast, Matthias Hagen, Benno Stein 0001 |
ICWSM | 1 |