Marcel Gohsen

dblp:204/0163 · DBLP profile ↗
← Back
12ranked-venue papers in the field
4as first author
11since 2021 · last 2026
0000-0002-1020-6745ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 11 (4 first)Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2026 Does Cognitive Load Affect Human Accuracy in Detecting Voice-Based Deepfakes?
abstract
Deepfake technologies are powerful tools that can be misused for malicious purposes such as spreading disinformation on social media. The effectiveness of such malicious applications depends on the ability of deepfakes to deceive their audience. Therefore, researchers have investigated human abilities to detect deepfakes in various studies. However, most of these studies were conducted with participants who focused exclusively on the detection task; hence the studies may not provide a complete picture of human abilities to detect deepfakes under realistic conditions: Social media users are exposed to cognitive load on the platform, which can impair their detection abilities. In this paper, we investigate the influence of cognitive load on human detection abilities of voice-based deepfakes in an empirical study with 30 participants. Our results suggest that low cognitive load does not generally impair detection abilities, and that the simultaneous exposure to a secondary stimulus can actually benefit people in the detection task.
Marcel Gohsen, Nicola Lea Libera, Johannes Kiesel, Jan Ehlers 0001, Benno Stein 0001
CHIIR1
2026 TREC iKAT 2025: A Test Collection for the Offline and Interactive Evaluation of Conversational Search
abstract
Conversational search agents, especially with the advent of large language models, have developed into useful tools to satisfy complex information needs of their users. Former research has shown that personalization (i.e., adaptation of agent responses to the preferences and traits of the user) can increase the relevance and perceived answer quality of these systems even further. However, developing accurate personalization methods typically requires rich datasets, both in terms of user profiles and complex conversations, for which only a few resources are publicly available. Over the past three years, the goal of the TREC Interactive Knowledge Assistance Track (iKAT) has been to bridge this gap. In organizing this shared task, we have developed a collection of complex information needs and associated conversations, made to challenge today's conversational agents and thus highlight aspects in need of further research. In this paper, we present the resources made for iKAT 2025, focusing on multi-session conversations (i.e., multiple dialogues per user), dynamically evolving user models, mixed-initiative dialogues, and large-scale human and automatic assessments. In addition to manually designed user profiles and conversations, the test collection for 2025 also contains dialogues between participating systems and our user simulators. All the resources are publicly available in our repository, including the system evaluation both as a result of the offline (i.e., test collection-based) and interactive tasks (i.e., user simulation-based), as well as their source code and model weights, to foster future research in this direction.
Zahra Abbasiantaeb, Simon Lupart, Marcel Gohsen, Nailia Mirzakhmedova, Johannes Kiesel, Jeff Dalton 0001, Mohammad Aliannejadi
SIGIR3
2026 The 10th Workshop on Search-Oriented Conversational Artificial Intelligence (SCAI'26)
abstract
SCAI (https://scai.info) celebrates its 10th anniversary this year and we would like to invite our core research community to join us. Since our first workshop started back at ICTIR 2017 in Amsterdam, we came a long way and would like to use this opportunity to reflect on it together. With the advent of large language models, conversational AI has emerged as a primary paradigm for search-intensive tasks. However, despite the vast success of conversational AI, there are major shortcomings in existing solutions that offer promising opportunities for the next breakthroughs which we would like to promote further. The focus of this edition will be on the personalization of conversational search systems, with a featured session for the former TREC shared task "Interactive Knowledge Assistance Track" (iKAT) reintroduced this year at SCAI. In combination with a panel discussion, invited presentations and keynote talks from major industry representatives, a lively poster session, and a separate break-out session featuring hands-on evaluation of the top-notch conversational AI systems, we plan for a full-day dense and highly engaging workshop.
Philipp Christmann, Roxana Petcu, Sneha Singhania, Mohammad Aliannejadi, Marcel Gohsen, Svitlana Vakulenko
SIGIR5
2026 Sim.API: A Middleware to Simplify the Use of User Simulators for Shared Tasks in Conversational Search
Marcel Gohsen, Nailia Mirzakhmedova, Zahra Abbasiantaeb, Johannes Kiesel, Simon Lupart, Jeff Dalton 0001, Benno Stein 0001, Mohammad Aliannejadi
SIGIR1
2026 Do Simulated Users Need to Remember? Analyzing the Impact of Memory Models in Conversational Search Evaluation
abstract
Conversational search systems are typically evaluated using a fixed reference collection of conversations or through user studies with a live system. However, fixed-reference conversations can cover only a few plausible conversations, and user studies are costly, time-consuming, and often hard to reproduce. A promising alternative that avoids coverage and cost issues is user simulation, in which a computer program takes on the role of a user and interacts with the system under evaluation. But the complexity of human search behavior raises the question of how ''realistic'' the simulations actually need to be for reliable evaluations of conversational search systems. In this paper, we ask: Do simulated users need to remember? While real users may learn and forget information during conversational search sessions, which inspired previous research to also model memory capabilities in simulations, it remains unclear whether this actually influences the results of system evaluations. To investigate the impact of memory modeling, we analyze conversations of simulated users and of humans with four conversational search systems. Our results suggest that incorporating long-term memory into simulators can help reproduce system effectiveness rankings obtained from human conversations, whereas incorporating short-term memory can diminish the reproduction. We also find that simulators are generally valid and reproducible---and memory modeling even increases run-to-run reproducibility of system rankings---but overall, simulations approximate human evaluation scores better when ''helpful'' assistants are evaluated than when assistants with deteriorated response quality are assessed. Our code and data are available at https://github.com/webis-de/SIGIR-26.
Nailia Mirzakhmedova, Marcel Gohsen, Johannes Kiesel, Matthias Hagen, Benno Stein 0001
SIGIR2
2025 Overview of Touché 2025: Argumentation Systems - Extended Abstract
Johannes Kiesel, Çagri Çöltekin, Marcel Gohsen, Sebastian Heineking, Maximilian Heinrich, Maik Fröbe, Tim Hagen, Mohammad Aliannejadi, Tomaz Erjavec, Matthias Hagen, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Harrisen Scells, Ines Zelch, Martin Potthast, Benno Stein 0001
ECIR (5)3
2024 Assisted Knowledge Graph Authoring: Human-Supervised Knowledge Graph Construction from Natural Language
abstract
Encyclopedic knowledge graphs, such as Wikidata, host an extensive repository of millions of knowledge statements. However, domain-specific knowledge from fields such as history, physics, or medicine is significantly underrepresented in those graphs. Although few domain-specific knowledge graphs exist (e.g., Pubmed for medicine), developing specialized retrieval applications for many domains still requires constructing knowledge graphs from scratch. To facilitate knowledge graph construction, we introduce WAKA: a Web application that allows domain experts to create knowledge graphs through the medium with which they are most familiar: natural language.
Marcel Gohsen, Benno Stein 0001
CHIIR1
2024 Simulating Follow-Up Questions in Conversational Search
Johannes Kiesel, Marcel Gohsen, Nailia Mirzakhmedova, Matthias Hagen, Benno Stein 0001
ECIR (2)2
2023 Guiding Oral Conversations: How to Nudge Users Towards Asking Questions?
abstract
How could an envisioned voice-based conversational information system assist the information seeker when the seeker does not know how to continue the conversation? The system could explicitly suggest a question to ask after each of its responses, but this approach quickly feels restrictive, repetitive, and interrupts immersion in the conversation. In this paper, we explore, for the first time, unobtrusive syntactic and auditive modifications of oral system responses to nudge information seekers towards asking about specific topics. We report the results of a crowdsourcing study with 965 participations that investigated the effectiveness and drawbacks of different modifications in three information scenarios.
Marcel Gohsen, Johannes Kiesel, Mariam Korashi, Jan Ehlers 0001, Benno Stein 0001
CHIIR1
2022 What is That? Crowdsourcing Questions to a Virtual Exhibition
abstract
Virtual environments with an ambient natural interface that allows to retrieve information for learning about the environment are a promising combination for implementing engaging virtual exhibitions. As a step towards better understanding search behavior in such exhibitions, this paper contributes the data from an exploratory study with participants asking questions on a real-world historical room, the Gropiuszimmer at the Weimar Bauhaus, while being on an “online virtual tour” through the room. The dataset comprises 849 manually categorized questions (557 in English, 292 in German) from 63 participants combined with a detailed interaction log, which allows replaying each session (29 hours total). The presented dataset and analyses aim to provide researchers and practitioners with a starting point to develop in-depth studies and prototypical systems.
Johannes Kiesel, Volker Bernhard, Marcel Gohsen, Josef Roth, Benno Stein 0001
CHIIR3
2022 Query Interpretations from Entity-Linked Segmentations
abstract
Web search queries can be ambiguous: is "source of the nile'' meant to find information on the actual river or on a board game of that name? We tackle this problem by deriving entity-based query interpretations: given some query, the task is to derive all reasonable ways of linking suitable parts of the query to semantically compatible entities in a background knowledge base. Our suggested approach focuses on effectiveness but also on efficiency since web search response times should not exceed some hundreds of milliseconds. In our approach, we use query segmentation as a pre-processing step that finds promising segment-based "interpretation skeletons''. The individual segments from these skeletons are then linked to entities from a knowledge base and the reasonable combinations are ranked in a final step. An experimental comparison on a combined corpus of all existing query entity linking datasets shows our approach to have a better interpretation accuracy at a better run time than the previously most effective methods.
Vaibhav Kasturia, Marcel Gohsen, Matthias Hagen
WSDM2
2017 A Large-Scale Query Spelling Correction Corpus
abstract
We present a new large-scale collection of 54,772 queries with manually annotated spelling corrections. For 9,170 of the queries (16.74%), spelling variants that are different to the original query are proposed. With its size, our new corpus is an order of magnitude larger than other publicly available query spelling corpora. In addition to releasing the new large-scale corpus, we also provide an implementation of the winner of the Microsoft Speller Challenge from~2011 and compare it on the different publicly available corpora to spelling corrections mined from Google and Bing. This way, we also shed some light on the spelling correction performance of state-of-the-art commercial search systems.
Matthias Hagen, Martin Potthast, Marcel Gohsen, Anja Rathgeber, Benno Stein 0001
SIGIR3