EDBT 2026 Demo / reviewers in the wild / expert
Preetam Prabhu Srikar Dammu
dblp:309/4843
· DBLP profile ↗
8ranked-venue papers
7as first author
8since 2021 · last 2026
0009-0007-2021-0354ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ClaimDB: A Fact Verification Benchmark over Large Structured DataabstractReal-world fact-checking often involves verifying claims grounded in structured data at scale.Despite substantial progress in factverification benchmarks, this setting remains largely underexplored.In this work, we introduce CLAIMDB, a fact-verification benchmark where the evidence for claims is derived from compositions of millions of records and multiple tables.CLAIMDB consists of 80 unique real-life databases covering a wide range of domains, from governance and healthcare to media, education and the natural sciences.At this scale, verification approaches that rely on "reading" the evidence break down, forcing a timely shift toward reasoning in executable programs.We conduct extensive experiments with 30 state-of-the-art proprietary and open-source (below 70B) LLMs and find that more than half score below 55% accuracy.Our analysis also reveals that both closed-and opensource models struggle with abstention-the ability to admit that there is no evidence to decide-raising doubts about their reliability in high-stakes data analysis tasks.We release the benchmark, code, and the LLM leaderboard at https://claimdb.github.io.BIRD Benchmark 10,962 6,545 Filter using SQL's AST . . .80 Unique DBs Execute SQL on DBs Entailed Prompt Contr. Michael Theologitis, Preetam Prabhu Srikar Dammu, Chirag Shah 0001, Dan Suciu |
ACL (1) | 2 |
| 2026 | Information Seeking in the Age of Agentic AI: A Half-Day TutorialabstractAgentic AI systems are changing how people seek and use information. Yet, the research community has not fully adapted – methods for studying, building, and assessing such systems often remain static, missing the interactive, temporal, and evidence-driven dynamics that characterize real information seeking. This hands-on tutorial equips the CHIIR community with a concise, practice-oriented methodology for leveraging and evaluating information-seeking agents. We define a shared vocabulary for agents and connect it to user-centered IR constructs; we show how to design agentic workflows that elicit effective evidence seeking under temporal change (planning, tool choice, grounding); and we introduce log-based rubrics that score correctness, evidence support, adequacy, and cost. Short case studies and optional demonstrations using open frameworks (for example, Perplexica, local LLMs via Ollama, and metasearch engines such as SearXNG) illustrate how these ideas map to real systems. Attendees receive reusable materials, including slides and selected supplemental resources (e.g., example traces and optional demo notebooks), suitable for research and teaching. The tutorial assumes familiarity with core IR concepts but does not require prior experience with agent frameworks. Preetam Prabhu Srikar Dammu, Tanya G. Roosta |
CHIIR | 1 |
| 2026 | iAgentBench: Benchmarking Sensemaking Capabilities of Information-Seeking Agents on High-Traffic TopicsabstractWith the emergence of search-enabled generative QA systems, users are increasingly turning to tools that browse, aggregate, and reconcile evidence across multiple sources on their behalf. Yet many widely used QA benchmarks remain answerable by retrieving a single relevant passage, making them poorly suited for measuring cross-source sensemaking, such as integrating evidence, tracking causal links, and resolving dependencies across facets of a topic. We present iAgentBench, a dynamic ODQA benchmark that targets these higher-level information needs while keeping questions natural and grounded in realistic information-seeking behavior. iAgentBench draws seed topics from real-world attention signals and uses common user intent patterns to construct user-like questions whose answers require combining evidence from multiple sources, not just extracting a single snippet. Each instance is released with traceable evidence and auditable intermediate artifacts that support contamination checks and enable fine-grained diagnosis of failures in retrieval versus synthesis. Experiments across multiple LLMs show that retrieval improves accuracy, but retrieval alone does not reliably resolve these questions, underscoring the need to evaluate evidence use, not just evidence access. Preetam Prabhu Srikar Dammu, Arnav Palkhiwala, Tanya G. Roosta, Chirag Shah 0001 |
SIGIR | 1 |
| 2025 | Dynamic-KGQA: A Scalable Framework for Generating Adaptive Question Answering DatasetsabstractAs question answering (QA) systems advance alongside the rapid evolution of foundation models, the need for robust, adaptable, and large-scale evaluation benchmarks becomes increasingly critical. Traditional QA benchmarks are often static and publicly available, making them susceptible to data contamination and memorization by large language models (LLMs). Consequently, static benchmarks may overestimate model generalization and hinder a reliable assessment of real-world performance. In this work, we introduce Dynamic-KGQA, a scalable framework for generating adaptive QA datasets from knowledge graphs (KGs), designed to mitigate memorization risks while maintaining statistical consistency across iterations. Unlike fixed benchmarks, Dynamic-KGQA generates a new dataset variant on every run while preserving the underlying distribution, enabling fair and reproducible evaluations. Furthermore, our framework provides fine-grained control over dataset characteristics, supporting domain-specific and topic-focused QA dataset generation. Additionally, Dynamic-KGQA produces compact, semantically coherent subgraphs that facilitate both training and evaluation of KGQA models, enhancing their ability to leverage structured knowledge effectively. To align with existing evaluation protocols, we also provide static large-scale train/test/validation splits, ensuring comparability with prior methods. By introducing a dynamic, customizable benchmarking paradigm, Dynamic-KGQA enables a more rigorous and adaptable evaluation of QA systems. Preetam Prabhu Srikar Dammu, Himanshu Naidu, Chirag Shah 0001 |
SIGIR | 1 |
| 2025 | Towards Ethical and Personalized Web Navigation Agents: A Framework for User-Aligned Task ExecutionabstractGenerative AI has advanced the capabilities of autonomous agents, enabling autonomous execution of complex web navigation tasks that can reshape digital interactions across various domains. Yet, to reach their full potential, these agents must be ethically aligned and personalized to individual user needs-a challenge complicated by privacy concerns and the risk of reinforcing biases. This work introduces a novel framework that enables responsible, user-guided personalization of web navigation agents, ensuring alignment with ethical standards and user preferences. By developing agents capable of perceiving, reasoning, and adapting in alignment with user preferences, this work proposes an approach that transcends generic task execution. Employing a structured representation of user-specific tasks, the agent utilizes interactive and reasoning actions to personalize workflows, adapting responsively to individual contexts. Evaluation through task success metrics and user satisfaction scores further assesses the ethical alignment and utility of personalized interactions. This research lays the groundwork for responsible agents that offer personalized assistance while adhering to ethical and privacy standards, with implications for information retrieval, e-commerce, and other knowledge-intensive applications. Preetam Prabhu Srikar Dammu |
WSDM | 1 |
| 2025 | A Shopping Agent for Addressing Subjective Product NeedsabstractIn e-commerce, customers often struggle to find relevant items when their needs involve subjective properties characterized by personal or collective perception, tastes, and opinions, which are typically not captured in catalog data. This challenge is particularly pronounced in event-based scenarios like gifting, where selecting the right product involves complex subjective reasoning. Customer reviews can be a valuable source of subjective information to bridge this gap. Consequently, customers often spend significant amount of time navigating multiple products and reading numerous reviews to find suitable gifts that meet their needs. In order to reduce the effort involved, we propose an agentic approach driven by large language models to streamline this process by autonomously executing various user actions. These include computational tasks like vagueness detection and subjective product needs extraction, conversational interactions to gather missing user information, and web browsing actions that search for product details, reviews, and review images. Additionally, the agent employs generative actions to synthesize gifting ideas and explanations, helping users discover suitable products more efficiently. The proposed approach not only reduces the cognitive burden on users but also facilitates the exploration of a wider range of products. Our solution highlights the potential of autonomous agents to handle subjective queries in e-commerce, enhancing personalization, product exploration, and selection in a user-centric manner. Preetam Prabhu Srikar Dammu, Omar Alonso, Barbara Poblete |
WSDM | 1 |
| 2024 | "They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated ConversationsabstractLarge language models (LLMs) have emerged as an integral part of modern societies, powering user-facing applications such as personal assistants and enterprise applications like recruitment tools.Despite their utility, research indicates that LLMs perpetuate systemic biases.Yet, prior works on LLM harms predominantly focus on Western concepts like race and gender, often overlooking cultural concepts from other parts of the world.Additionally, these studies typically investigate "harm" as a singular dimension, ignoring the various and subtle forms in which harms manifest.To address this gap, we introduce the Covert Harms and Social Threats (CHAST), a set of seven metrics grounded in social science literature.We utilize evaluation models aligned with human assessments to examine the presence of covert harms in LLM-generated conversations, particularly in the context of recruitment.Our experiments reveal that seven out of the eight LLMs included in this study generated conversations riddled with CHAST, characterized by malign views expressed in seemingly neutral language unlikely to be detected by existing methods.Notably, these LLMs manifested more extreme views and opinions when dealing with non-Western concepts like caste, compared to Western ones such as race.Warning: This paper has instances of offensive language to serve as examples. Preetam Prabhu Srikar Dammu, Hayoung Jung, Monojit Choudhury, Tanushree Mitra |
EMNLP | 1 |
| 2023 | Addressing Weak Decision Boundaries in Image Classification by Leveraging Web Search and Generative ModelsabstractMachine learning (ML) technologies are known to be riddled with ethical and operational problems, however, we are witnessing an increasing thrust by businesses to deploy them in sensitive applications. One major issue among many is that ML models do not perform equally well for underrepresented groups. This puts vulnerable populations in an even disadvantaged and unfavorable position. We propose an approach that leverages the power of web search and generative models to alleviate some of the shortcomings of discriminative models. We demonstrate our method on an image classification problem using ImageNet's People Subtree subset, and show that it is effective in enhancing robustness and mitigating bias in certain classes that represent vulnerable populations (e.g., female doctor of color). Our new method is able to (1) identify weak decision boundaries for such classes; (2) construct search queries for Google as well as text for generating images through DALL-E 2 and Stable Diffusion; and (3) show how these newly captured training samples could alleviate population bias issue. While still improving the model's overall performance considerably, we achieve a significant reduction (77.30%) in the model's gender accuracy disparity. In addition to these improvements, we observed a notable enhancement in the classifier's decision boundary, as it is characterized by fewer weakspots and an increased separation between classes. Although we showcase our method on vulnerable populations in this study, the proposed technique is extendable to a wide range of problems and domains. Preetam Prabhu Srikar Dammu, Yunhe Feng, Chirag Shah 0001 |
IJCAI | 1 |