VLDB 2026 Research / reviewers in the wild / expert
Watheq Mansour
dblp:300/0963 · also Watheq Ahmad Mansour
· DBLP profile ↗
10ranked-venue papers in the field
8as first author
10since 2021 · last 2026
0000-0002-9463-595XORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 10 (8 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Audio Retrieval: Challenges and Solutions
Watheq Mansour |
ECIR (3) | 1 |
| 2026 | Revisiting Human-vs-LLM Judgments Using the TREC Podcast Track
Watheq Mansour, J. Shane Culpepper, Joel Mackenzie, Andrew Yates |
ECIR (2) | 1 |
| 2026 | How Variability Influences Podcast Search: Queries, Transcriptions, and JudgesabstractPodcasts have continued to grow in popularity over the last two decades, with more than 4.52 million podcasts and 584 million listeners across the globe in 2025. Developing effective search systems for web-scale podcast corpora is of vital importance. Previous research has approached the search task primarily through text representation using a single transcription of the audio content via automatic speech recognition (ASR) models. However, there is currently limited understanding about how variation in podcast representations influences ranking, retrieval, and relevance assessment. Watheq Mansour, J. Shane Culpepper, Andrew Yates, Joel Mackenzie |
SIGIR | 1 |
| 2025 | Examining the Impact of Transcript Variation on Podcast Search and Re-ranking
Watheq Mansour, J. Shane Culpepper, Joel Mackenzie |
ECIR (3) | 1 |
| 2025 | Towards Efficient and Effective Multi-modal RetrievalabstractInformation Retrieval systems are becoming increasingly multimodal, and one interesting and natural modality for human information interaction is audio. Audio data is rapidly expanding in both volume and popularity. For example, more than 4 million podcasts are available worldwide - each containing tens to thousands of episodes - with over 500 million expected listeners as of 2024. Searching such large and evolving audio corpora involves multiple challenges, ranging from content and style diversity, expensive audio processing, and various indexing considerations such as retrieval segmentation and text vs audio retrieval modalities. In this thesis, we aim to address the challenges linked to audio retrieval and devise novel approaches to improve state-of-the-art (SOTA) performance. We divide the work into four main milestones that target answering the following research questions: RQ1: How does the quality of an ASR model affect retrieval effectiveness? One of the approaches to search spoken audio is to reduce the problem to text retrieval using ASR models to generate text transcripts. Spina et al. [2] showed that the quality of ASR has a clear impact on the summaries generated by ASR. Motivated by this, we are interested in examining how automatic transcription algorithms impact the effectiveness of audio retrieval and re-ranking in dense and sparse retrieval configurations. In many real-world scenarios, resource limits require cheaper ASR models to be used. Therefore, we are interested in exploring if dense semantic search methods can reduce the impact of ASR model choice. RQ2: How can we build more complete, comprehensive, and representative audio retrieval resources to promote further research on audio retrieval? After reviewing the literature and performing a preliminary experiment on the Spotify podcasts dataset [1], we noted two main gaps in the available resources. Firstly, the large audio dataset released in TREC suffers from the shallow pooling issue, where new systems retrieve many unseen judgments, making it difficult to assess the quality of modern IR techniques. Secondly, there is a scarcity of audio datasets available for low-resource languages. Our plan to bridge these gaps comprises two parts: (1) Assessing the missed judgments Motivated by the recent literature on using LLMs as a judge [3], we plan to resolve the unjudged passages issue by assessing the relevance with multiple large language models (LLMs). (2) Curate a multi-lingual audio dataset with a focus on low-resource languages. Since there is a scarcity of audio datasets in low-resource languages, we plan to collect a large number of audio files in multiple low-resource languages, unify their format, and make them publicly available as a first step toward building test collections that enable research and advancements in such languages. RQ3: How can we combine audio and text to produce an efficient and effective multi-modal ranking model? We will focus on solutions that depend on processing both text and audio signals during indexing and retrieval. Our first study will be focused on improving the effectiveness of low-cost ASR systems. Another promising direction is to use the phonetic features as a representation of a word or token. Then, we can devise a suitable indexing paradigm to support searching over these representations. RQ4: What are the challenges affecting audio indexing and search performance, and what is the best way to mitigate against similar issues in the future? One of the main challenges in multi-modal search is identifying the suitable length for the retrieval unit. The retrieval unit satisfying an information need might vary in length from a couple of minutes to 10 minutes to one hour or a whole episode or series. Other challenges will stem from failure cases within state-of-the-art text/audio retrieval systems. So, in parallel with answering the previous questions, we will also study such failure cases, investigate the reasons behind them, and propose solutions that can mitigate their negative impacts. Watheq Mansour |
SIGIR | 1 |
| 2024 | Revisiting Document Expansion and Filtering for Effective First-Stage RetrievalabstractDocument expansion is a technique that aims to reduce the likelihood of term mismatch by augmenting documents with related terms or queries. Doc2Query minus minus (Doc2Query-) represents an extension to the expansion process that uses a neural model to identify and remove expansions that may not be relevant to the given document, thereby increasing the quality of the ranking while simultaneously reducing the amount of augmented data. In this work, we conduct a detailed reproducibility study of Doc2Query- to better understand the trade-offs inherent to document expansion and filtering mechanisms. After successfully reproducing the best-performing method from the Doc2Query- family, we show that filtering actually harms recall-based metrics on various test collections. Next, we explore whether the two-stage "generate-then-filter" process can be replaced with a single generation phase via reinforcement learning. Finally, we extend our experimentation to learned sparse retrieval models and demonstrate that filtering is not helpful when term weights can be learned. Overall, our work provides a deeper understanding of the behaviour and characteristics of common document expansion mechanisms, and paves the way for developing more efficient yet effective augmentation models. Watheq Mansour, Shengyao Zhuang, Guido Zuccon, Joel Mackenzie |
SIGIR | 1 |
| 2023 | Tahaqqaq: A Real-Time System for Assisting Twitter Users in Arabic Claim VerificationabstractOver the past years, notable progress has been made towards fighting misinformation spread over social media, encouraging the development of many fact-checking systems. However, systems that operate over Arabic content are scarce. In this work, we bridge this gap by proposing Tahaqqaq (Verify), an Arabic real-time system that helps users verify claims over Twitter with several functionalities, such as identifying check-worthy claims, estimating credibility of users in terms of spreading fake news, and finding authoritative accounts. Tahaqqaq has a friendly online Web interface that supports various real-time user scenarios. In the same breath, we enable public access to Tahaqqaq services through a handy RESTful API. Finally, in terms of performance, multiple components of Tahaqqaq outperform the state-of-the-art models on Arabic datasets. Zien Sheikh Ali, Watheq Mansour, Fatima Haouari, Maram Hasanain, Tamer Elsayed, Abdulaziz Alali 0001 |
SIGIR | 2 |
| 2023 | Who can verify this? Finding authorities for rumor verification in TwitterabstractA large body of research work has proposed verification techniques for rumors spreading in social media that mainly relied on subjective evidence, e.g., propagation networks or user interactions. Alternatively, in this work, we introduce the task of authority finding in social media, in which we aim to find authorities, for given rumors spreading specifically in Twitter, who can help verify them by providing exclusive/convincing evidence that supports or denies those rumors. We release the first test collection for Authority FINding in Arabic Twitter (AuFIN). The collection comprises 150 rumors (expressed in tweets) associated with a total of 1,044 authority accounts and a user collection of 395,231 Twitter accounts (members of 1,192,284 unique Twitter lists). Moreover, we propose a hybrid model that employs pre-trained language models and combines lexical, semantic, and network signals to find authorities. Our experiments show that the textual representation of users is insufficient, and incorporating the Twitter network features improved the recall of authorities by 34%. Moreover, semantic ranking is inferior to the lexical and network-based ranking in terms of precision, but superior in terms of recall. Therefore, combining both the semantic and network-based ranking achieved the best overall performance achieving a precision of 0.413 and 0.213 at depth 1 and 5 respectively. We show that rumor expansion by exploiting Knowledge Bases improves the recall of authorities by up to 15%. Furthermore, we find that SOTA models for topic expert finding perform poorly on finding authorities. Finally, drawing upon our experiments, we discuss failure factors and make recommendations for future research directions in addressing this task. Fatima Haouari, Tamer Elsayed, Watheq Mansour |
Inf. Process. Manag. | 3 |
| 2023 | This is not new! Spotting previously-verified claims over TwitterabstractSeveral fake claims are commonly repeated over time, especially on social media. To identify such previous claims, the verified claim retrieval task was studied, where, for a given input claim, the goal is to find previously-verified claims that are relevant to it. However, this view assumes that each claim was already verified, which may not be true for all claims in the real-world scenario. In this work, we introduce the Verified Claim Checking problem over Twitter, in which the relevant verified claims are retrieved only if the input claim was indeed previously-verified, thus saving computation time. We address the problem by proposing SpotVC, an end-to-end approach consisting of two stages, namely a filter and a reranker. The proposed filter achieved an average F1 of 0.81 while significantly reducing computation time. Moreover, the proposed reranker outperformed the state-of-the-art models on two public datasets and provided on-par performance on a third one. Overall, our proposed system exhibits an effective operational balance in the trade-off between efficiency and effectiveness for the real-world scenario. Watheq Mansour, Tamer Elsayed, Abdulaziz Alali 0001 |
Inf. Process. Manag. | 1 |
| 2022 | Did I See It Before? Detecting Previously-Checked Claims over Twitter
Watheq Mansour, Tamer Elsayed, Abdulaziz Alali 0001 |
ECIR (1) | 1 |