VLDB 2026 Research / reviewers in the wild / expert
Danish Contractor
dblp:93/9012
· DBLP profile ↗
25ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0002-6843-1961ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reducing the Scope of Language ModelsabstractLarge language models (LLMs) are deployed in a wide variety of user-facing applications. Typically, these deployments have some specific purpose, like answering questions grounded on documentation or acting as coding assistants, but they require general language understanding. In such deployments, LLMs should respond only to queries that align with the intended purpose and reject all other requests, such as generating poetry or answering questions about physics, a task we refer to as 'scoping'. We conduct a comprehensive empirical evaluation of various methods, ranging from prompting, fine-tuning to preference learning and the recently proposed general alignment technique known as Circuit Breakers (CB). Across three families of language models and a broad variety of tasks, we show that it is possible to scope language models. We examine scoping for multiple topics, and fine-grained topics. We ablate diversity of irrelevant queries, layer different techniques, conduct adversarial evaluations and more. Among other results, we find that when diverse examples of irrelevant queries are available, simple supervised fine-tuning produces the best results, but when such diversity is low, Circuit Breakers perform quite well. One can often get the benefits of both methods by layering them in succession. We intend our study to serve as a practitioner's guide to scoping LLMs. David Yunis, Siyu Huo, R. Chulaka Gunasekara, Danish Contractor |
AAAI | 4 |
| 2026 | KCIF: Knowledge-Conditioned Instruction Following
Rudra Murthy, Praveen Venkateswaran, Danish Contractor |
LREC | 4 |
| 2025 | mtRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation SystemsabstractAbstract Retrieval-augmented generation (RAG) has recently become a very popular task for Large Language Models (LLMs). Evaluating them on multi-turn RAG conversations, where the system is asked to generate a response to a question in the context of a preceding conversation, is an important and often overlooked task with several additional challenges. We present mtRAG, an end-to-end human-generated multi-turn RAG benchmark that reflects several real-world properties across diverse dimensions for evaluating the full RAG pipeline. mtRAG contains 110 conversations averaging 7.7 turns each across four domains for a total of 842 tasks. We also explore automation paths via synthetic data and LLM-as-a-Judge evaluation. Our human and automatic evaluations show that even state-of-the-art LLM RAG systems struggle on mtRAG. We demonstrate the need for strong retrieval and generation systems that can handle later turns, unanswerable questions, non-standalone questions, and multiple domains. mtRAG is available at https://github.com/ibm/mt-rag-benchmark. Yannis Katsis, Sara Rosenthal, Kshitij Fadnis, R. Chulaka Gunasekara, Young-Suk Lee 0001, Lucian Popa 0001, Vraj Shah, Huaiyu Zhu 0001, Danish Contractor, Marina Danilevsky |
Trans. Assoc. Comput. Linguistics | 9 |
| 2024 | Position: Standardization of Behavioral Use Clauses is Necessary for the Adoption of Responsible Licensing of AIabstractGrowing concerns over negligent or malicious uses of AI have increased the appetite for tools that help manage the risks of the technology. In 2018, licenses with behaviorial-use clauses (commonly referred to as Responsible AI Licenses) were proposed to give developers a framework for releasing AI assets while specifying their users to mitigate negative applications. As of the end of 2023, on the order of 40,000 software and model repositories have adopted responsible AI licenses licenses. Notable models licensed with behavioral use clauses include BLOOM (language) and LLaMA2 (language), Stable Diffusion (image), and GRID (robotics). This paper explores why and how these licenses have been adopted, and why and how they have been adapted to fit particular use cases. We use a mixed-methods methodology of qualitative interviews, clustering of license clauses, and quantitative analysis of license adoption. Based on this evidence we take the position that responsible AI licenses need standardization to avoid confusing users or diluting their impact. At the same time, customization of behavioral restrictions is also appropriate in some contexts (e.g., medical domains). We advocate for “standardized customization” that can meet users’ needs and can be supported via tooling. Daniel McDuff, Tim Korjakow, Scott Cambo, Jesse Josua Benjamin, Jenny Lee, Yacine Jernite, Carlos Muñoz Ferrandis, Aaron Gokaslan, Alek Tarkowski, Joseph Lindley, A. Feder Cooper, Danish Contractor |
ICML | 12 |
| 2023 | Prompting with Pseudo-Code InstructionsabstractPrompting with natural language instructions has recently emerged as a popular method of harnessing the capabilities of large language models (LLM).Given the inherent ambiguity present in natural language, it is intuitive to consider the possible advantages of prompting with less ambiguous prompt styles, like pseudocode.In this paper, we explore if prompting via pseudo-code instructions helps improve the performance of pre-trained language models.We manually create a dataset 1 of pseudo-code prompts for 132 different tasks spanning classification, QA, and generative language tasks, sourced from the Super-NaturalInstructions dataset (Wang et al., 2022b).Using these prompts along with their counterparts in natural language, we study their performance on two LLM families -BLOOM (Scao et al., 2023), CodeGen (Nijkamp et al., 2023).Our experiments show that using pseudo-code instructions leads to better results, with an average increase (absolute) of 7-16 points in F1 scores for classification tasks and an improvement (relative) of 12-38% in aggregate ROUGE-L scores across all tasks.We include detailed ablation studies which indicate that code comments, docstrings, and the structural clues encoded in pseudo-code all contribute towards the improvement in performance.To the best of our knowledge, our work is the first to demonstrate how pseudocode prompts can be helpful in improving the performance of pre-trained LMs.* Equal contribution 1 Code and dataset available at https://github.com/ mayank31398/pseudo-code-instructions Listing 1 An example pseudo-code instruction for the task from Wang et al. (2022b).A successful model is expected to use the provided pseudo-code instructions and output responses to a pool of evaluation instances.1 def generate_sentiment(sentence: str) -> str: 2 """For the given sentence, the task is to 3 predict the sentiment.For positive 4 sentiment return "positive" else return 5 "negative". Riyaz A. Bhat, Rudra Murthy V, Danish Contractor, Srikanth Tamilselvam |
EMNLP | 5 |
| 2022 | Variational Learning for Unsupervised Knowledge Grounded DialogsabstractRecent methods for knowledge grounded dialogs generate responses by incorporating information from an external textual document. These methods do not require the exact document to be known during training and rely on the use of a retrieval system to fetch relevant documents from a large index. The documents used to generate the responses are modeled as latent variables whose prior probabilities need to be estimated. Models such as RAG, marginalize the document probabilities over the documents retrieved from the index to define the log-likelihood loss function which is optimized end-to-end. In this paper, we develop a variational approach to the above technique wherein, we instead maximize the Evidence Lower bound (ELBO). Using a collection of three publicly available open-conversation datasets, we demonstrate how the posterior distribution, which has information from the ground-truth response, allows for a better approximation of the objective function during training. To overcome the challenges associated with sampling over a large knowledge collection, we develop an efficient approach to approximate the ELBO. To the best of our knowledge, we are the first to apply variational training for open-scale unsupervised knowledge grounded dialog systems. Dhiraj Madan, Gaurav Pandey 0001, Danish Contractor |
IJCAI | 4 |
| 2021 | Bootstrapping Dialog Models from Human to Human Conversation LogsabstractState-of-the-art commercial dialog platforms provide powerful tools to build a conversational agent. These platforms provide complete control to the dialog designer to model user-agent interactions. However, a dialog designer needs to rely on domain experts to manually build the dialog model -- by creating dialog flow nodes and modeling user intents. This process is laborious, time consuming and expensive and does not allow the designer to exploit human to human conversation logs effectively. In this work, we present a research prototype that can ingest human-to-human conversation logs between an end-user and an agent, and suggest user-intents and agent-responses, given a conversation context. We utilize human to human conversation logs to build two emulators: user and agent. An agent emulator models an agent response given the conversation context so far, and a user emulator outputs possible user responses. Our system is able to recommend conversational intents as well as conversation flow using emulators based on real-world data, thus making the process of designing a bot more efficient. To the best our knowledge this is the first system that enables data-driven dialog model creation by emulating users and agents. Pankaj Dhoolia, Danish Contractor, Sachindra Joshi |
AAAI | 3 |
| 2021 | Answering POI-recommendation Questions using Tourism ReviewsabstractWe introduce the novel and challenging task of answering Points-of-interest (POI) recommendation questions, using a collection of reviews that describe candidate answer entities (POIs). We harvest a QA dataset that contains 47,124 paragraph-sized user questions from travelers seeking POI recommendations for hotels, attractions and restaurants. Each question can have thousands of candidate entities to choose from and each candidate is associated with a collection of unstructured reviews. Questions can include requirements based on physical location, budget, timings as well as other subjective considerations related to ambience, quality of service etc. Our dataset requires reasoning over a large number of candidate answer entities (over 5300 per question on average) and we find that running commonly used neural architectures for QA is prohibitively expensive. Further, commonly used retriever-ranker based methods also do not work well for our task due to the nature of review-documents. Thus, as a first attempt at addressing some of the novel challenges of reasoning-at-scale posed by our task, we present a task specific baseline model that uses a three-stage cluster-select-rerank architecture. The model first clusters text for each entity to identify exemplar sentences describing an entity. It then uses a neural information retrieval (IR) module to select a set of potential entities from the large candidate set. A reranker uses a deeper attention-based architecture to pick the best answers from the selected entities. This strategy performs better than a pure retrieval or a pure attention-based reasoning approach yielding nearly 25% relative improvement in [email protected] over both approaches. To the best of our knowledge we are the first to present an unstructured QA-style task for POI-recommendation, using real-world tourism questions and POI-reviews. Danish Contractor, Krunal Shah 0001, Aditi Partap, Parag Singla, Mausam |
CIKM | 1 |
| 2021 | Joint Spatio-Textual Reasoning for Answering Tourism QuestionsabstractOur goal is to answer real-world tourism questions that seek Points-of-Interest (POI) recommendations. Such questions express various kinds of spatial and non-spatial constraints, necessitating a combination of textual and spatial reasoning. In response, we develop the first joint spatio-textual reasoning model, which combines geo-spatial knowledge with information in textual corpora to answer questions. We first develop a modular spatial-reasoning network that uses geo-coordinates of location names mentioned in a question, and of candidate answer POIs, to reason over only spatial constraints. We then combine our spatial-reasoner with a textual reasoner in a joint model and present experiments on a real world POI recommendation task. We report substantial improvements over existing models without joint spatio-textual reasoning. To the best of our knowledge, we are the first to develop a joint QA model that combines reasoning over external geo-spatial knowledge along with textual reasoning. Danish Contractor, Shashank Goel, Mausam, Parag Singla |
WWW | 1 |
| 2021 | Constrained BERT BiLSTM CRF for understanding multi-sentence entity-seeking questionsabstractAbstract We present the novel task of understanding multi-sentenceentity-seekingquestions (MSEQs), that is, the questions that may be expressed in multiple sentences, and that expect one or more entities as an answer. We formulate the problem of understanding MSEQs as a semantic labeling task over an open representation that makes minimal assumptions about schema or ontology-specific semantic vocabulary. At the core of our model, we use a BiLSTM (bidirectional LSTM) conditional random field (CRF), and to overcome the challenges of operating with low training data, we supplement it by using BERT embeddings, hand-designed features, as well as hard and soft constraints spanning multiple sentences. We find that this results in a 12–15 points gain over a vanilla BiLSTM CRF. We demonstrate the strengths of our work using the novel task of answering real-world entity-seeking questions from the tourism domain. The use of our labels helps answer 36% more questions with 35% more (relative) accuracy as compared to baselines. We also demonstrate how our framework can rapidly enable the parsing of MSEQs in an entirely new domain with small amounts of training data and little change in the semantic representation. Danish Contractor, Barun Patra, Mausam, Parag Singla |
Nat. Lang. Eng. | 1 |
| 2020 | Neural Conversational QA: Learning to Reason vs Exploiting PatternsabstractNeural Conversational QA tasks like ShARC require systems to answer questions based on the contents of a given passage.On studying recent state-of-the-art models on the ShARC QA task, we found indications that the models learn spurious clues/patterns in the dataset.Furthermore, we show that a heuristic-based program designed to exploit these patterns can have performance comparable to that of the neural models.In this paper we share our findings about four types of patterns found in the ShARC corpus and describe how neural models exploit them.Motivated by the aforementioned findings, we create and share a modified dataset that has fewer spurious patterns, consequently allowing models to learn better. Nikhil Verma, Dhiraj Madan, Danish Contractor, Sachindra Joshi |
EMNLP (1) | 4 |
| 2018 | Exemplar Encoder-Decoder for Neural Conversation GenerationabstractIn this paper we present the Exemplar Encoder-Decoder network (EED), a novel conversation model that learns to utilize similar examples from training data to generate responses.Similar conversation examples (context-response pairs) from training data are retrieved using a traditional TF-IDF based retrieval model.The retrieved responses are used to create exemplar vectors that are used by the decoder to generate the response.The contribution of each retrieved response is weighed by the similarity of corresponding context with the input context.We present detailed experiments on two large data sets and find that our method outperforms state of the art sequence to sequence generative models on several recently proposed evaluation metrics.We also observe that the responses generated by the proposed EED model are more informative and diverse compared to existing state-of-the-art method. Gaurav Pandey 0001, Danish Contractor, Sachindra Joshi |
ACL (1) | 2 |
| 2018 | Document Chunking and Learning Objective Generation for Instruction Design
Khoi-Nguyen Tran, Jey Han Lau, Danish Contractor, Bikram Sengupta, Christopher J. Butler, Mukesh K. Mohania |
EDM | 3 |
| 2016 | Document Segmentation for Labeling with Academic Learning Objectives
Divyanshu Bhartiya, Danish Contractor, Sovan Biswas, Bikram Sengupta, Mukesh K. Mohania |
EDM | 2 |
| 2016 | Entity-balanced Gaussian pLSA for Automated Comparison
Danish Contractor, Parag Singla, Mausam |
HLT-NAACL | 1 |
| 2015 | Tracking Political Elections on Social Media: Applications and Experience
Danish Contractor, Bhupesh Chawda, Sameep Mehta, L. Venkata Subramaniam, Tanveer A. Faruquie |
IJCAI | 1 |
| 2015 | Labeling Educational Content with Academic Learning StandardsabstractLearning standards (frequently referred to as academic standards, course curriculum etc.) define the specific structure of an educational program. Learning standards contain a list of instructions specifying various skills that students should learn at different points during their learning progression. For example,“calculate the area of a triangle” is one such instruction in a 6th grade geometry curriculum. Currently these instructions are imparted using prescribed textbooks or lesson plans which have been labeled with learning standard instructions. Teachers and students use this labeled learning content to identify relevant material for teaching and studying. However with an increasing amount of users as well as publisher generated content in recent days, teachers and students may want to refer to additional content apart from prescribed textbooks for their teaching/learning needs which is not labeled with learning standard instructions. Manually identifying the appropriate learning standard instruction for each learning content is time consuming and not scalable especially since learning standards frequently contain thousands of instructions, and subject to periodic revision. In this paper, we address the problem of automatically labeling digital learning content with the learning standards. Towards this goal, we first build semantic representations of the learning standard instructions using external knowledge sources such as Wikipedia and domain text books. These semantic representations are then used in a framework which utilizes structural constraints imposed by the hierarchy of the learning standards to assign labels to the learning materials. We demonstrate the usefulness of our approach on a collection of high school learning materials that were labeled by curriculum experts from a US school district according to a publicly available learning standard. The system developed has been deployed and is in use by the school district. To the best of our knowledge we are the first to attempt this novel task and develop such a system. Danish Contractor, Kashyap Popat, Shajith Ikbal, Sumit Negi, Bikram Sengupta, Mukesh K. Mohania |
SDM | 1 |
| 2013 | Content Analytics System for Social Customer Relationship Management
Meena Nagarajan, Danish Contractor, Stephen Dill, Jitendra Ajmera, Hyung-Il Ahn, Ashish Verma 0001, Matthew Denesuk |
ICWSM | 2 |
| 2013 | A CRM system for social media: challenges and experiencesabstractThe social Customer Relationship Management (CRM) landscape is attracting significant attention from customers and enterprises alike as a sustainable channel for tracking, managing and improving customer relations. Enterprises are taking a hard look at this open, unmediated platform because the community effect generated on this channel can have a telling effect on their brand image, potential market opportunity and customer loyalty. In this work we present our experiences in building a system that mines conversations on social platforms to identify and prioritize those posts and messages that are relevant to enterprises. The system presented in this work aims to empower an agent or a representative in an enterprise to monitor, track and respond to customer communication while also encouraging community participation. Jitendra Ajmera, Hyung-Il Ahn, Meena Nagarajan, Ashish Verma 0001, Danish Contractor, Stephen Dill, Matthew Denesuk |
WWW | 5 |
| 2012 | Using Argumentative Zones for Extractive Summarization of Scientific Articles
Danish Contractor, Anna Korhonen |
COLING | 1 |
| 2012 | Using content and interactions for discovering communities in social networksabstractIn recent years, social networking sites have not only enabled people to connect with each other using social links but have also allowed them to share, communicate and interact over diverse geographical regions. Social network provide a rich source of heterogeneous data which can be exploited to discover previously unknown relationships and interests among groups of people. In this paper, we address the problem of discovering topically meaningful communities from a social network. We assume that a persons' membership in a community is conditioned on its social relationship, the type of interaction and the information communicated with other members of that community. We propose generative models that can discover communities based on the discussed topics, interaction types and the social connections among people. In our models a person can belong to multiple communities and a community can participate in multiple topics. This allows us to discover both community interests and user interests based on the information and linked associations. We demonstrate the effectiveness of our model on two real word data sets and show that it performs better than existing community discovery models. Mrinmaya Sachan, Danish Contractor, Tanveer A. Faruquie, L. Venkata Subramaniam |
WWW | 2 |
| 2011 | Probabilistic model for discovering topic based communities in social networksabstractSocial graphs have received renewed interest as a research topic with the advent of social networking websites. These online networks provide a rich source of data to study user relationships and interaction patterns on a large scale. In this paper, we propose a generative Bayesian model for extracting latent communities from a social graph. We assume that community memberships depend on topics of interest between users and the link relationships between them in the social graph topology. In addition, we make use of the nature of interaction to gauge user interests. Our model allows communities to be related to multiple topics and each user in the graph can be a member of multiple communities. This gives an insight into user interests and topical distribution in communities. We show the effectiveness of our model using a real world data set and also compare our model with existing community discovery methods. Mrinmaya Sachan, Danish Contractor, Tanveer A. Faruquie, L. Venkata Subramaniam |
CIKM | 2 |
| 2011 | Labeling Unlabeled Data using Cross-Language Guided Clustering
Sachindra Joshi, Danish Contractor, Sumit Negi |
IJCNLP | 2 |
| 2011 | Auto-Grouping Emails For Faster E-Discovery
Sachindra Joshi, Danish Contractor, Kenney Ng, Prasad Deshpande, Thomas Hampp |
Proc. VLDB Endow. | 2 |
| 2010 | Handling Noisy Queries in Cross Language FAQ Retrieval
Danish Contractor, Govind Kothari, Tanveer A. Faruquie, L. Venkata Subramaniam, Sumit Negi |
EMNLP | 1 |