VLDB 2026 Research / reviewers in the wild / expert
Zafaryab Rasool
dblp:258/3376
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-3603-3125ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RAGProbe: Breaking RAG Pipelines with Evaluation ScenariosabstractRetrieval Augmented Generation (RAG) is increasingly employed in building Generative AI applications, yet their evaluation often relies on manual, trial-and-error processes. Automating this evaluation process involves generating test data to trigger failures involving context comprehension, data formatting, specificity, and content completeness. Random question-answer generation is insufficient. However, prior works rely on standard QA datasets, benchmarks and tactics that are not tailored to the specific domain requirements. Hence, current approaches and datasets do not trigger sufficiently broad and context-specific failures. In this paper, we introduce evaluation scenarios that describe the process of generating question-answer pairs from content indexed by RAG pipelines, and they are designed to trigger a wider range of failures and to simplify automation. This enables developers to identify and address weaknesses more effectively. We validate our approach on five open-source RAG pipelines using three datasets. Our approach triggers high failure rates, by generating prompts that combine multiple questions (up to 91% failure rate) highlighting the need for developers to prioritize handling such queries. We generated failure rates of 60% in an academic domain dataset and 53% and 64% in open-domain datasets. Compared to existing state-of-the-art methods, our approach triggers 77% more failures on average per RAG pipeline and 53% more failures on average per dataset, offering a mechanism to support developers to improve the RAG pipeline quality. Shangeetha Sivasothy, Scott Barnett, Stefanus Kurniawan, Zafaryab Rasool, Rajesh Vasa |
CAIN | 4 |
| 2024 | LLMs for Test Input Generation for Semantic ApplicationsabstractLarge language models (LLMs) enable state-of-the-art semantic capabilities to be added to software systems such as semantic search of unstructured documents and text generation. However, these models are computationally expensive. At scale, the cost of serving thousands of users increases massively affecting also user experience. To address this problem, semantic caches are used to check for answers to similar queries (that may have been phrased differently) without hitting the LLM service. Due to the nature of these semantic cache techniques that rely on query embeddings, there is a high chance of errors impacting user confidence in the system. Adopting semantic cache techniques usually requires testing the effectiveness of a semantic cache (accurate cache hits and misses) which requires a labelled test set of similar queries and responses which is often unavailable. In this paper, we present VaryGen, an approach for using LLMs for test input generation that produces similar questions from unstructured text documents. Our novel approach uses the reasoning capabilities of LLMs to 1) adapt queries to the domain, 2) synthesise subtle variations to queries, and 3) evaluate the synthesised test dataset. We evaluated our approach in the domain of a student question and answer system by qualitatively analysing 100 generated queries and result pairs, and conducting an empirical case study with an open source semantic cache. Our results show that query pairs satisfy human expectations of similarity and our generated data demonstrates failure cases of a semantic cache. Additionally, we also evaluate our approach on Qasper dataset. This work is an important first step into test input generation for semantic applications and presents considerations for practitioners when calibrating a semantic cache. Zafaryab Rasool, Scott Barnett, David Willie, Stefanus Kurniawan, Sherwin Balugo, Srikanth Thudumu, Mohamed Almorsy |
CAIN | 1 |
| 2023 | Data-dependent and Scale-Invariant Kernel for Support Vector Machine Classification
Vinayaka Vivekananda Malgi, Sunil Aryal, Zafaryab Rasool, David Tay |
PAKDD (1) | 3 |
| 2023 | Overcoming weaknesses of density peak clustering using a data-dependent similarity measure
Zafaryab Rasool, Sunil Aryal, Mohamed Reda Bouadjenek, Richard Dazeley |
Pattern Recognit. | 1 |
| 2022 | Index-Based Solutions for Efficient Density Peak ClusteringabstractDensity Peak Clustering (DPC), a popular density-based clustering approach, has received considerable attention from the research community primarily due to its simplicity and fewer-parameter requirement. However, the resultant clusters obtained using DPC are influenced by the sensitive parameter$d_c$, which depends on data distribution and requirements of different users. Besides, the original DPC algorithm requires visiting a large number of objects, making it slow. To this end, this paper investigates index-based solutions for DPC. Specifically, we propose two list-based index methods viz. (i) a simple List Index, and (ii) an advanced Cumulative Histogram Index. Efficient query algorithms are proposed for these indices which significantly avoids irrelevant comparisons at the cost of space. For memory-constrained systems, we further introduce an approximate solution to the above indices which allows substantial reduction in the space cost, provided that slight inaccuracies are admissible. Furthermore, owing to considerably lower memory requirements of existing tree-based index structures, we also present effective pruning techniques and efficient query algorithms to support DPC using the popular Quadtree Index and R-tree Index. Finally, we practically evaluate all the above indices and present the findings and results, obtained from a set of extensive experiments on six synthetic and real datasets. The experimental insights obtained can help to guide in selecting a befitting index. Zafaryab Rasool, Rui Zhou 0001, Lu Chen 0008, Chengfei Liu, Jiajie Xu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Index-based Solutions for Efficient Density Peak Clustering (Extended Abstract)abstractClusters reflect a potential relationship among different entities of data. This data can be sourced from a wide range of domains like market research, spatial data analysis, etc. Many clustering algorithms have been developed in the last few decades in response to the proliferating demands across industries and organizations, which help them make operational and strategic decisions. Among them, density-based clustering algorithms are popular, which find subsets of objects in "dense regions" separated by not-so-dense regions, where each subset represents a cluster. In this paper, our focal point will be Density Peak Clustering (DPC) [1] , a popular approach towards obtaining density-based clusters. Zafaryab Rasool, Rui Zhou 0001, Lu Chen 0008, Chengfei Liu, Jiajie Xu 0001 |
ICDE | 1 |