VLDB 2026 Research / reviewers in the wild / expert
Nailia Mirzakhmedova
dblp:339/1001
· DBLP profile ↗
6ranked-venue papers in the field
1as first author
6since 2021 · last 2026
0000-0002-8143-1405ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 6 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TREC iKAT 2025: A Test Collection for the Offline and Interactive Evaluation of Conversational SearchabstractConversational search agents, especially with the advent of large language models, have developed into useful tools to satisfy complex information needs of their users. Former research has shown that personalization (i.e., adaptation of agent responses to the preferences and traits of the user) can increase the relevance and perceived answer quality of these systems even further. However, developing accurate personalization methods typically requires rich datasets, both in terms of user profiles and complex conversations, for which only a few resources are publicly available. Over the past three years, the goal of the TREC Interactive Knowledge Assistance Track (iKAT) has been to bridge this gap. In organizing this shared task, we have developed a collection of complex information needs and associated conversations, made to challenge today's conversational agents and thus highlight aspects in need of further research. In this paper, we present the resources made for iKAT 2025, focusing on multi-session conversations (i.e., multiple dialogues per user), dynamically evolving user models, mixed-initiative dialogues, and large-scale human and automatic assessments. In addition to manually designed user profiles and conversations, the test collection for 2025 also contains dialogues between participating systems and our user simulators. All the resources are publicly available in our repository, including the system evaluation both as a result of the offline (i.e., test collection-based) and interactive tasks (i.e., user simulation-based), as well as their source code and model weights, to foster future research in this direction. Zahra Abbasiantaeb, Simon Lupart, Marcel Gohsen, Nailia Mirzakhmedova, Johannes Kiesel, Jeff Dalton 0001, Mohammad Aliannejadi |
SIGIR | 4 |
| 2026 | Sim.API: A Middleware to Simplify the Use of User Simulators for Shared Tasks in Conversational Search
Marcel Gohsen, Nailia Mirzakhmedova, Zahra Abbasiantaeb, Johannes Kiesel, Simon Lupart, Jeff Dalton 0001, Benno Stein 0001, Mohammad Aliannejadi |
SIGIR | 2 |
| 2026 | Do Simulated Users Need to Remember? Analyzing the Impact of Memory Models in Conversational Search EvaluationabstractConversational search systems are typically evaluated using a fixed reference collection of conversations or through user studies with a live system. However, fixed-reference conversations can cover only a few plausible conversations, and user studies are costly, time-consuming, and often hard to reproduce. A promising alternative that avoids coverage and cost issues is user simulation, in which a computer program takes on the role of a user and interacts with the system under evaluation. But the complexity of human search behavior raises the question of how ''realistic'' the simulations actually need to be for reliable evaluations of conversational search systems. In this paper, we ask: Do simulated users need to remember? While real users may learn and forget information during conversational search sessions, which inspired previous research to also model memory capabilities in simulations, it remains unclear whether this actually influences the results of system evaluations. To investigate the impact of memory modeling, we analyze conversations of simulated users and of humans with four conversational search systems. Our results suggest that incorporating long-term memory into simulators can help reproduce system effectiveness rankings obtained from human conversations, whereas incorporating short-term memory can diminish the reproduction. We also find that simulators are generally valid and reproducible---and memory modeling even increases run-to-run reproducibility of system rankings---but overall, simulations approximate human evaluation scores better when ''helpful'' assistants are evaluated than when assistants with deteriorated response quality are assessed. Our code and data are available at https://github.com/webis-de/SIGIR-26. Nailia Mirzakhmedova, Marcel Gohsen, Johannes Kiesel, Matthias Hagen, Benno Stein 0001 |
SIGIR | 1 |
| 2025 | Overview of Touché 2025: Argumentation Systems - Extended Abstract
Johannes Kiesel, Çagri Çöltekin, Marcel Gohsen, Sebastian Heineking, Maximilian Heinrich, Maik Fröbe, Tim Hagen, Mohammad Aliannejadi, Tomaz Erjavec, Matthias Hagen, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Harrisen Scells, Ines Zelch, Martin Potthast, Benno Stein 0001 |
ECIR (5) | 14 |
| 2024 | Overview of Touché 2024: Argumentation Systems
Johannes Kiesel, Çagri Çöltekin, Maximilian Heinrich, Maik Fröbe, Milad Alshomary, Bertrand De Longueville, Tomaz Erjavec, Nicolas Handke, Matyás Kopp, Nikola Ljubesic, Katja Meden, Nailia Mirzakhmedova, Vaidas Morkevicius, Theresa Reitis-Münstermann, Mario Scharfbillig, Nicolas Stefanovitch, Henning Wachsmuth, Martin Potthast, Benno Stein 0001 |
ECIR (5) | 12 |
| 2024 | Simulating Follow-Up Questions in Conversational Search
Johannes Kiesel, Marcel Gohsen, Nailia Mirzakhmedova, Matthias Hagen, Benno Stein 0001 |
ECIR (2) | 3 |