Leila Tavakoli

dblp:164/1008 · DBLP profile ↗
← Back
8ranked-venue papers in the field
5as first author
7since 2021 · last 2026
0000-0002-5951-4052ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 8 (5 first)
YearPublicationVenuePosition
2026 Policy-Guided RAG: Enforcing Verbatim and Controlled Synthesis
Mina Naghash Asadi, Leila Tavakoli, Mustafa Bilgrami
SIGIR2
2026 Deletion Isn't Enough: Auditing RAG for Selective Forgetting
Leila Tavakoli, Mark Sanderson
SIGIR1
2025 Online and Offline Evaluation in Search Clarification
abstract
The effectiveness of clarification question models in engaging users within search systems is currently constrained, casting doubt on their overall usefulness. To improve the performance of these models, it is crucial to employ assessment approaches that encompass both real-time feedback from users (online evaluation) and the characteristics of clarification questions evaluated through human assessment (offline evaluation). However, the relationship between online and offline evaluations has been debated in information retrieval. This study aims to investigate how this discordance holds in search clarification. We use user engagement as ground truth and employ several offline labels to investigate to what extent the offline ranked lists of clarification resemble the ideal ranked lists based on online user engagement. Contrary to the current understanding that offline evaluations fall short of supporting online evaluations, we indicate that when identifying the most engaging clarification questions from the user’s perspective, online and offline evaluations correspond with each other. We show that the query length does not influence the relationship between online and offline evaluations, and reducing uncertainty in online evaluation strengthens this relationship. We illustrate that an engaging clarification needs to excel from multiple perspectives, and SERP quality and characteristics of the clarification are equally important. We also investigate if human labels can enhance the performance of Large Language Models (LLMs) and Learning-to-Rank (LTR) models in identifying the most engaging clarification questions from the user’s perspective by incorporating offline evaluations as input features. Our results indicate that LTR models do not perform better than individual offline labels. However, GPT, an LLM, emerges as the standout performer, surpassing all LTR models and offline labels.
Leila Tavakoli, Johanne R. Trippas, Hamed Zamani, Falk Scholer, Mark Sanderson
ACM Trans. Inf. Syst.1
2022 MIMICS-Duo: Offline & Online Evaluation of Search Clarification
abstract
Asking clarification questions is an active area of research; however, resources for training and evaluating search clarification methods are not sufficient. To address this issue, we describe MIMICS-Duo, a new freely available dataset of 306 search queries with multiple clarifications (a total of 1,034 query-clarification pairs). MIMICS-Duo contains fine-grained annotations on clarification questions and their candidate answers and enhances the existing MIMICS datasets by enabling multi-dimensional evaluation of search clarification methods, including online and offline evaluation. We conduct extensive analysis to demonstrate the relationship between offline and online search clarification datasets and outline several research directions enabled by MIMICS-Duo. We believe that this resource will help researchers better understand clarification in search.
Leila Tavakoli, Johanne R. Trippas, Hamed Zamani, Falk Scholer, Mark Sanderson
SIGIR1
2022 Analyzing clarification in asynchronous information-seeking conversations
abstract
Abstract This research analyzes human‐generated clarification questions to provide insights into how they are used to disambiguate and provide a better understanding of information needs. A set of clarification questions is extracted from posts on the Stack Exchange platform. Novel taxonomy is defined for the annotation of the questions and their responses. We investigate the clarification questions in terms of whether they add any information to the post (the initial question posted by the asker) and the accepted answer, which is the answer chosen by the asker. After identifying, which clarification questions are more useful, we investigated the characteristics of these questions in terms of their types and patterns. Non‐useful clarification questions are identified, and their patterns are compared with useful clarifications. Our analysis indicates that the most useful clarification questions have similar patterns, regardless of topic. This research contributes to an understanding of clarification in conversations and can provide insight for clarification dialogues in conversational search scenarios and for the possible system generation of clarification requests in information‐seeking conversations.
Leila Tavakoli, Hamed Zamani, Falk Scholer, W. Bruce Croft, Mark Sanderson
J. Assoc. Inf. Sci. Technol.1
2021 Quantifying Human-Perceived Answer Utility in Non-factoid Question Answering
abstract
Taking a user-centric approach, we study the features that render an answer to a non-factoid question useful in the eyes of the person who asked that question. An editorial study, where participants assess the usefulness of the answers they received in response to their questions, as well as 12 different aspects associated with the answers, indicates considerable correlation between certain aspects such as relevance, correctness, and completeness with the user-perceived usefulness of answers. Moreover, we investigate the effectiveness of some commonly used answer quality measures, such as ROGUE, BLEU, METEOR, and BERTScore, demonstrating that these measures are limited in their ability to capture the aspects of usefulness and have room for improvement. The question answering dataset created in our work was made publicly available.
Berkant Barla Cambazoglu, Valeria Bolotova-Baranova, Falk Scholer, Mark Sanderson, Leila Tavakoli, W. Bruce Croft
CHIIR5
2021 An Intent Taxonomy for Questions Asked in Web Search
abstract
We present a new, multi-faceted taxonomy to classify questions asked in web search engines based on the question intent, types of entities mentioned, types of question words, and granularity of the expected answer. Built based on the inspection of 1,000 real-life questions issued to a web search engine, the taxonomy reflects the recent search behavior of users and enables deep understanding of user intents, goals, and expected answers. This taxonomy is more fine-grained than previous query taxonomies, and is designed with the ultimate goal of reducing the inherent ambiguity in determining the intent of questions. In addition, we describe the formal procedure for conducting an editorial study of the taxonomy including its evaluation. The adopted procedure aims to increase assessor agreement without incurring too much overhead. Our results demonstrate that, despite being more fine-grained, the proposed intent categories result in higher agreement between assessors compared to an existing, commonly used taxonomy.
Berkant Barla Cambazoglu, Leila Tavakoli, Falk Scholer, Mark Sanderson, W. Bruce Croft
CHIIR2
2020 Generating Clarifying Questions in Conversational Search Systems
abstract
Asking a clarifying question can be a key element improving the performance of information seeking systems, particularly conversational search systems due to their limited bandwidth interfaces. While generating and asking clarifying questions is important; get-ting an answer for the clarifying question is also essential as a clarifying question without an answer is useless. Therefore, as the first step in current research, we analysed human-generated clarifying questions in a Community Question Answering website as a sample of conversation. This helped us to gain a better insight into how users interact with clarification. We investigated the clarifying questions in terms of whether they add any information to the question and the accepted answer. We further discovered the patterns and types of such clarifying questions. The next phase of this research will then generate clarifying questions in conversational search systems. We will then employ neural network models to generate clarifying questions to maximise clarification in questions. The proposed model will be trained using the MIMICS data collection in addition to our collected dataset. We will also attempt to consider the recognised patterns from the analysis conducted in the first step to enhance the chance that a user will interact with the clarifying questions. Finally, we will aim to minimise the interaction between the search system and the user to reduce the risk of dropping the conversation by the user due to asking too many clarifying questions.
Leila Tavakoli
CIKM1