Heydar Soudani

dblp:320/7388 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0003-0393-8662ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Total Recall QA: A Verifiable Evaluation Suite for Deep Research Agents
abstract
Deep research agents have emerged as LLM-based systems designed to perform multi-step information seeking and reasoning over large, open-domain sources to answer complex questions by synthesizing information from multiple information sources. Given the complexity of the task and despite various recent efforts, evaluation of deep research agents remains fundamentally challenging. This paper identifies a list of requirements and optional properties for evaluating deep research agents. We observe that existing benchmarks do not satisfy all identified requirements. Inspired by prior research on TREC Total Recall Tracks, we introduce the task of Total Recall Question Answering and develop a framework for deep research agents evaluation that satisfies the identified criteria. Our framework constructs single-answer, total recall queries with precise evaluation and relevance judgments derived from a structured knowledge base paired with a text corpus, enabling large-scale data construction. Using this framework, we build TRQA, a deep research benchmark constructed from Wikidata-Wikipedia as a real-world source and a synthetically generated e-commerce knowledge base and corpus to mitigate the effects of data contamination. We benchmark the collection with representative retriever and deep research models and establish baseline retrieval and end-to-end results for future comparative evaluation.
Mahta Rafiee, Heydar Soudani, Zahra Abbasiantaeb, Mohammad Aliannejadi, Faegheh Hasibi, Hamed Zamani
SIGIR2
2026 Uncertainty Quantification for Retrieval-Augmented Reasoning
abstract
Retrieval-augmented reasoning (RAR) is a recent evolution of retrieval-augmented generation (RAG) that employs multiple reasoning steps for retrieval and generation. While effective for some complex queries, RAR remains vulnerable to errors and misleading outputs. Uncertainty quantification (UQ) offers methods to estimate the confidence of systems' outputs. These methods, however, often handle simple queries with no retrieval or single-step retrieval, without properly handling RAR setup. Accurate estimation of UQ for RAR requires accounting for all sources of uncertainty, including those arising from retrieval and generation. In this paper, we account for these sources and introduce Retrieval-Augmented Reasoning Consistency (R2C), a novel UQ method for RAR. The core idea of R2C is to perturb the multi-step reasoning process by applying various actions to reasoning steps. These perturbations alter the retriever's input, which shifts its output and consequently modifies the generator's input at the next step. Through this iterative feedback loop, the retriever and generator continuously reshape each other's inputs, enabling us to capture uncertainty arising from both components. Experiments on five popular RAR systems across diverse QA datasets show that R2C improves AUROC by over 5% on average compared to the state-of-the-art UQ baselines. Extrinsic evaluations using R2C as an external signal further confirm its effectiveness for two downstream tasks: in the Abstention task, it achieves ~5% gains in both F1Abstain and AccAbstain; in Model Selection, it improves exact match by ~7% over single models and ~3% over selection methods. Code is available on https://github.com/HeydarSoudani/R2C.
Heydar Soudani, Hamed Zamani, Faegheh Hasibi
SIGIR1
2025 Enhancing Knowledge Injection in Large Language Models for Efficient and Trustworthy Responses
abstract
Large Language Models (LLMs) have shown remarkable proficiency in Natural Language Generation (NLG) across various tasks. However, they often require additional resources beyond their internal knowledge to respond reliably to user queries. Determining the optimal methods, content, and timing for introducing new knowledge remains a critical challenge without clear solutions. Our main objective is to enhance knowledge injection in LLMs to generate trustworthy responses. Therefore, this research is centered around two main research questions. (RQ1) What are the most effective and efficient choices of knowledge injection for question-answering over less popular knowledge? A key challenge in knowledge injection is determining how to introduce new knowledge into an LLM while balancing both effectiveness and efficiency. RAG and FT with synthetic data have emerged as two distinct paradigms, yet there has been no comprehensive comparison that highlights both their strengths and limitations. To address this, we conduct an extensive evaluation of RAG and FT for handling less popular factual knowledge, assuming limited textual descriptions are available for a given domain and application. Through this analysis, we find that RAG substantially outperforms FT in this setup. Our second research question is: (RQ2 ) How can we quantify the uncertainty of LLMs during response generation and leverage it to improve the reliability of their outputs? By knowing LLMs uncertainty, one can determine when to leverage RAG and when to rely solely on LLMs' internal knowledge. However, existing Uncertainty Estimation (UE) techniques primarily focus on scenarios where the input consists only of a user query, overlooking the complexities introduced by retrieved knowledge. We investigate UE in the context of RAG and find that the performance of current UE methods is inconsistent, often degrading when non-parametric knowledge is incorporated into the input prompt. We propose an axiomatic framework to formalize optimal behavior of UE methods. Recent advancements in active RAG aim to enhance the dynamic interaction between retrievers and generators. In this context, UE is primarily employed to determine when retrieval is necessary. Typically, uncertainty is measured based on next-token probabilities, and the interpretation of an uncertainty value is based on comparing it with accuracy. However, we raise three key questions: (RQ2.1) What is the optimal method for measuring uncertainty? We argue that relying solely on probability-based methods may not be the most effective approach. Furthermore, in conversational systems, previous turns or sessions do not directly influence next-token probabilities, highlighting the need for a new UE method. (RQ2.2) How should uncertainty values be interpreted? Current approaches, which define uncertainty through relative comparisons with other values, lack precision. Additionally, some applications correlate uncertainty with correctness, but how should uncertainty be interpreted in cases where correctness labels are unavailable? Moreover, we aim to enhance conversational interactions by leveraging an appropriate UE method. Specifically, we will investigate the research question: (RQ2.3) Can an LLM anticipate its next action, identify its information needs, and then generate an appropriate response? To explore this, we will utilize UE to help the LLM determine its next step, whether to generate a response, retrieve relevant information, ask a clarification question, or take an alternative action.
Heydar Soudani
SIGIR1
2023 Data Augmentation for Conversational AI
abstract
Advancements in conversational systems have revolutionized information access, surpassing the limitations of single queries. However, developing dialogue systems requires a large amount of training data, which is a challenge in low-resource domains and languages. Traditional data collection methods like crowd-sourcing are labor-intensive and time-consuming, making them ineffective in this context. Data augmentation (DA) is an affective approach to alleviate the data scarcity problem in conversational systems. This tutorial provides a comprehensive and up-to-date overview of DA approaches in the context of conversational systems. It highlights recent advances in conversation augmentation, open domain and task-oriented conversation generation, and different paradigms of evaluating these models. We also discuss current challenges and future directions in order to help researchers and practitioners to further advance the field in this area.
Heydar Soudani, Evangelos Kanoulas, Faegheh Hasibi
CIKM1
2022 Persian Natural Language Inference: A Meta-learning Approach
abstract
Incorporating information from other languages can improve the results of tasks in low-resource languages. A powerful method of building functional natural language processing systems for low-resource languages is to combine multilingual pre-trained representations with cross-lingual transfer learning. In general, however, shared representations are learned separately, either across tasks or across languages. This paper proposes a meta-learning approach for inferring natural language in Persian. Alternately, meta-learning uses different task information (such as QA in Persian) or other language information (such as natural language inference in English). Also, we investigate the role of task augmentation strategy for forming additional high-quality tasks. We evaluate the proposed method using four languages and an auxiliary task. Compared to the baseline approach, the proposed model consistently outperforms it, improving accuracy by roughly six percent. We also examine the effect of finding appropriate initial parameters using zero-shot evaluation and CCA similarity.
Heydar Soudani, Mohammad Hassan Mojab, Hamid Beigy
COLING1