VLDB 2026 Research / reviewers in the wild / expert
Mausam
dblp:30/6391
· DBLP profile ↗
99ranked-venue papers
10as first author
34since 2021 · last 2026
0000-0003-4088-4296ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 92 · 10 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 6 first-author · 3 since 2021Databases, data management, data science and information retrieval · 11 · 3 since 2021Human-computer interaction and ubiquitous computing · 6Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Combining Distantly Supervised Models with In Context Learning for Monolingual and Cross-Lingual Relation ExtractionabstractDistantly Supervised Relation Extraction (DSRE) remains a long-standing challenge in NLP, where models must learn from noisy bag-level annotations while making sentencelevel predictions.While existing state-of-theart (SoTA) DSRE models rely on task-specific training, their integration with in-context learning (ICL) using large language models (LLMs) remains underexplored.A key challenge is that the LLM may not learn relation semantics correctly, due to noisy annotation.In response, we propose HYDRE -HYbrid Distantly Supervised Relation Extraction framework.It first uses a trained DSRE model to identify the top-k candidate relations for a given test sentence, then uses a novel dynamic exemplar retrieval strategy that extracts reliable, sentence-level exemplars from training data, which are then provided in LLM prompt for outputting the final relation(s).We further extend HYDRE to cross-lingual settings for RE in low-resource languages.Using available English DSRE training data, we evaluate all methods on English as well as a newly curated benchmark covering four diverse low-resource Indic languages -Oriya, Santali, Manipuri, and Tulu.HYDRE achieves up to 20 F1 point gains in English and, on average, 17 F1 points on Indic languages over prior SoTA DSRE models and naive prompting baselines.Detailed ablations exhibit HYDRE's efficacy compared to other prompting strategies. Vipul Rathore, Malik Hammad Faisal, Parag Singla, Mausam |
ACL (1) | 4 |
| 2025 | STARQA: A Question Answering Dataset for Complex Analytical Reasoning over Structured DatabasesabstractSemantic parsing methods for converting text to SQL queries enable question answering over structured data and can greatly benefit analysts who routinely perform complex analytics on vast data stored in specialized relational databases.Although several benchmarks measure the abilities of text to SQL, the complexity of their questions is inherently limited by the level of expressiveness in query languages and none focus explicitly on questions involving complex analytical reasoning which require operations such as calculations over aggregate analytics, time series analysis or scenario understanding.In this paper, we introduce STARQA, the first public humancreated dataset of complex analytical reasoning questions and answers on three specializeddomain databases.In addition to generating SQL directly using LLMs, we evaluate a novel approach (TEXT2SQLCODE) that decomposes the task into a combination of SQL and Python: SQL is responsible for data fetching, and Python more naturally performs reasoning.Our results demonstrate that identifying and combining the abilities of SQL and Python is beneficial compared to using SQL alone, yet the dataset still remains quite challenging for the existing state-of-the-art LLMs. Mounica Maddela, Lingjue Xie, Daniel Preotiuc-Pietro, Mausam |
EMNLP | 4 |
| 2025 | A Framework for Leveraging Partially-Labeled Data for Product Attribute-Value IdentificationabstractIn the e-commerce domain, the accurate extraction of attribute-value pairs (e.g., B rand : Apple ) from product titles and user search queries is crucial for enhancing search and recommendation systems. A major challenge with neural models for this task is the lack of high-quality training data, as the annotations for attribute-value pairs in the available datasets are often incomplete. To address this, we introduce G en T o C, a model designed for training directly with partially-labeled data, eliminating the necessity for a fully annotated dataset. G en T o C employs a marker-augmented generative model to identify potential attributes, followed by a token classification model that determines the associated values for each attribute. G en T o C outperforms existing state-of-the-art models, exhibiting upto 56.3% increase in the number of accurate extractions. Furthermore, we utilize G en T o C to regenerate the training dataset to expand attribute-value annotations. This bootstrapping substantially improves the data quality for training other standard NER models, which are typically faster but less capable in handling partially-labeled data, enabling them to achieve comparable performance to G en T o C. Our results demonstrate G en T o C's unique ability to learn from a limited set of partially-labeled data and improve the training of more efficient models, advancing the automated extraction of attribute-value pairs. Finally, our model has been successfully integrated into IndiaMART, India's largest B2B e-commerce platform, achieving a significant increase of 20.2% in the number of correctly identified attribute-value pairs over the existing deployed system while achieving a high precision of 89.5%. We have released the code for G en T o C model at https://github.com/KnowDisAI/GenToC. D. Subhalingam, Keshav Kolluru, Mausam, Saurabh Singal |
KDD (1) | 3 |
| 2024 | GOALNET: Interleaving Neural Goal Predicate Inference with Classical Planning for Generalization in Robot Instruction FollowingabstractOur goal is to enable a robot to learn how to sequence its actions to perform high-level tasks specified as natural language instructions, given successful demonstrations from a human partner. Our novel neuro-symbolic solution GOALNET builds an iterative two-step approach that interleaves (i) inferring next subgoal predicate implied by the language instruction, for a given world state, and (ii) synthesizing a feasible subgoal-reaching plan from that state. The agent executes the plan, and the two steps are repeated. GOALNET combines (i) learning, where dense representations are acquired for language instruction and the world state via a neural network prediction model, enabling generalization to novel settings and (ii) planning, where the cause-effect modeling by a classical planner eschews irrelevant predicates, facilitating multi-stage decision making in large domains. GOALNET obtains 78% improvement in the goal reaching rate in comparison to several state-of-the-art approaches on benchmark data with multi-stage instructions. Further, GOALNET can generalize to novel instructions for scenes with unseen objects. Source code available at https://github. com/reail-iitd/goalnet. Jigyasa Gupta, Shreya Sharma 0005, Shreshth Tuli, Rohan Paul, Mausam |
AAAI | 5 |
| 2024 | RetinaQA: A Robust Knowledge Base Question Answering Model for both Answerable and Unanswerable QuestionsabstractAn essential requirement for a real-world Knowledge Base Question Answering (KBQA) system is the ability to detect answerability of questions when generating logical forms.However, state-of-the-art KBQA models assume all questions to be answerable.Recent research has found that such models, when superficially adapted to detect answerability, struggle to satisfactorily identify the different categories of unanswerable questions, and simultaneously preserve good performance for answerable questions.Towards addressing this issue, we propose RetinaQA, a new KBQA model that unifies two key ideas in a single KBQA architecture: (a) discrimination over candidate logical forms, rather than generating these, for handling schema-related unanswerability, and (b) sketch-filling-based construction of candidate logical forms for handling data-related unaswerability.Our results show that RetinaQA significantly outperforms adaptations of state-of-the-art KBQA models in handling both answerable and unanswerable questions and demonstrates robustness across all categories of unanswerability.Notably, Reti-naQA also sets a new state-of-the-art for answerable KBQA, surpassing existing models.We release our code-base 1 for further research. Prayushi Faldu, Indrajit Bhattacharya, Mausam |
ACL (1) | 3 |
| 2024 | Few-shot Transfer Learning for Knowledge Base Question Answering: Fusing Supervised Models with In-Context LearningabstractMayur Patidar, Riya Sawhney, Avinash Singh, Biswajit Chatterjee, Mausam ., Indrajit Bhattacharya. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Mayur Patidar, Riya Sawhney, Avinash Kumar Singh, Biswajit Chatterjee, Mausam, Indrajit Bhattacharya |
ACL (1) | 5 |
| 2024 | Synergizing In-context Learning with Hints for End-to-end Task-oriented Dialog SystemsabstractEnd-to-end Task-Oriented Dialog (TOD) systems typically require extensive training datasets to perform well.In contrast, large language model (LLM) based TOD systems can excel even with limited data due to their ability to learn tasks through in-context exemplars.However, these models lack alignment with the style of responses in training data and often generate comprehensive responses, making it difficult for users to grasp the information quickly.In response, we propose SyncTOD that synergizes LLMs with task-specific hints to improve alignment in low-data settings.Sync-TOD employs small auxiliary models to provide hints and select exemplars for in-context prompts.With ChatGPT, SyncTOD achieves superior performance compared to LLM-based baselines and SoTA models in low-data settings, while retaining competitive performance in full-data settings. Vishal Vivek Saley, Rocktim Jyoti Das, Dinesh Raghu, Mausam |
EMNLP | 4 |
| 2024 | MediTOD: An English Dialogue Dataset for Medical History Taking with Comprehensive AnnotationsabstractMedical task-oriented dialogue systems can assist doctors by collecting patient medical history, aiding in diagnosis, or guiding treatment selection, thereby reducing doctor burnout and expanding access to medical services.However, doctor-patient dialogue datasets are not readily available, primarily due to privacy regulations.Moreover, existing datasets lack comprehensive annotations involving medical slots and their different attributes, such as symptoms and their onset, progression, and severity.These comprehensive annotations are crucial for accurate diagnosis.Finally, most existing datasets are non-English, limiting their utility for the larger research community.In response, we introduce MediTOD, a new dataset of doctor-patient dialogues in English for the medical history-taking task.Collaborating with doctors, we devise a questionnairebased labeling scheme tailored to the medical domain.Then, medical professionals create the dataset with high-quality comprehensive annotations, capturing medical slots and their attributes.We establish benchmarks in supervised and few-shot settings on MediTOD for natural language understanding, policy learning, and natural language generation subtasks, evaluating models from both TOD and biomedical domains.We release MediTOD resources for future research.* Work done when authors were at IIT Delhi.[{"intent": "salutations"}] Doctor: How may I help you?[{"intent": "inform", "slots": { "positive_symptom": [ {"value": "pharyngitis", "onset": "past four days"}, {"value": "fever", "onset": "last two days"} ]}}] Patient: Yes, I just came in here today.I I've just been.Really getting like the soreness in my throat for the past, I would say four days and I also had a fever for the last two days as well. CMAS Key-Value Legends:[{"intent": "inform", "slots": { "positive_symptom": "pharyngitis", "positive_symptom": "fever", "onset": "past four days", "onset": "last two days" }] Datasets Language Annotations #utterances/ #utterancesAll TOD Tasks Comprehensive Canonicalized dialogue Vishal Vivek Saley, Goonjan Saha, Rocktim Jyoti Das, Dinesh Raghu, Mausam |
EMNLP | 5 |
| 2024 | AutoMix: Automatically Mixing Language ModelsabstractLarge language models (LLMs) are now available from cloud API providers in various sizes and configurations. While this diversity offers a broad spectrum of choices, effectively leveraging the options to optimize computational cost and performance remains challenging. In this work, we present AutoMix, an approach that strategically routes queries to larger LMs, based on the approximate correctness of outputs from a smaller LM. Central to AutoMix are two key technical contributions. First, it has a few-shot self-verification mechanism, which estimates the reliability of its own outputs without requiring extensive training. Second, given that self-verification can be noisy, it employs a POMDP based router that can effectively select an appropriately sized model, based on answer confidence. Experiments across five language models and five challenging datasets show that Automix consistently surpasses strong baselines, reducing computational cost by over 50\% for comparable performance. Pranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju, Swaroop Mishra, Aditya Gupta 0001, Dheeraj Rajagopal, Karthik Kappaganthu, Yiming Yang 0002, Shyam Upadhyay, Manaal Faruqui, Mausam |
NeurIPS | 13 |
| 2024 | Matching papers and reviewers at large conferencesabstractPeer-reviewed conferences, the main publication venues in CS, rely critically on matching highly qualified reviewers for each paper. Because of the growing scale of these conferences, the tight timelines on which they operate, and a recent surge in explicitly dishonest behavior, there is now no alternative to performing this matching in an automated way. This paper introduces Large Conference Matching (LCM), a novel reviewer–paper matching approach that was recently deployed in the 35th AAAI Conference on Artificial Intelligence (AAAI 2021), and has since been adopted (wholly or partially) by other conferences including ICML 2022, AAAI 2022-2024, and IJCAI 2022-2024. LCM has three main elements: (1) collecting and processing input data to identify problematic matches and generate reviewer–paper scores; (2) formulating and solving an optimization problem to find good reviewer–paper matchings; and (3) a two-phase reviewing process that shifts reviewing resources away from papers likely to be rejected and towards papers closer to the decision boundary. This paper also describes an evaluation of these innovations based on an extensive post-hoc analysis on real data—including a comparison with the matching algorithm used in AAAI's previous (2020) iteration—and supplements this with additional numerical experimentation.2 Kevin Leyton-Brown, Mausam, Yatin Nandwani, Hedayat Zarkoob, Chris Cameron, Neil Newman, Dinesh Raghu |
Artif. Intell. | 2 |
| 2023 | Multimodal Persona Based Generation of Comic DialogsabstractWe focus on the novel problem of persona based dialogue generation for comic strips.Dialogs in comic strips is a unique and unexplored area where every strip contains utterances from various characters with each one building upon the previous utterances and the associated visual scene.Previous works like Di-aloGPT, PersonaGPT and other dialog generation models encode two-party dialogues and do not account for the visual information.To the best of our knowledge we are the first to propose the paradigm of multimodal persona based dialogue generation.We contribute a novel dataset, COMSET, consisting of 54K strips, harvested from 13 popular comics available online.Further, we propose a multimodal persona-based architecture, MPDIALOG, to generate dialogues for the next panel in the strip which decreases the perplexity score by ∼10 points over strong dialogue generation baseline models.We demonstrate that there is still ample opportunity for improvement, highlighting the importance of building stronger dialogue systems that are able to generate persona-consistent dialogues and understand the context through various modalities. Harsh Agrawal, Manish Gupta 0001, Mausam |
ACL (1) | 4 |
| 2023 | DiSCoMaT: Distantly Supervised Composition Extraction from Tables in Materials Science ArticlesabstractTanishq Gupta, Mohd Zaki, Devanshi Khatsuriya, Kausik Hira, N M Anoop Krishnan, Mausam -. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Tanishq Gupta, Mohd Zaki, Devanshi Khatsuriya, Kausik Hira, N. M. Anoop Krishnan, Mausam |
ACL (1) | 6 |
| 2023 | Do I have the Knowledge to Answer? Investigating Answerability of Knowledge Base QuestionsabstractMayur Patidar, Prayushi Faldu, Avinash Singh, Lovekesh Vig, Indrajit Bhattacharya, Mausam -. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Mayur Patidar, Prayushi Faldu, Avinash Kumar Singh, Lovekesh Vig, Indrajit Bhattacharya, Mausam |
ACL (1) | 6 |
| 2023 | DeGPR: Deep Guided Posterior Regularization for Multi-Class Cell Detection and CountingabstractMulti-class cell detection and counting is an essential task for many pathological diagnoses. Manual counting is tedious and often leads to inter-observer variations among pathologists. While there exist multiple, general-purpose, deep learning-based object detection and counting methods, they may not readily transfer to detecting and counting cells in medical images, due to the limited data, presence of tiny overlapping objects, multiple cell types, severe class-imbalance, minute differences in size/shape of cells, etc. In response, we propose guided posterior regularization (DEGPR), which assists an object detector by guiding it to exploit discriminative features among cells. The features may be pathologist-provided or inferred directly from visual data. We validate our model on two publicly available datasets (CoNSeP and MoNuSAC), and on MuCeD, a novel dataset that we contribute. MuCeD consists of 55 biopsy images of the human duodenum for predicting celiac disease. We perform extensive experimentation with three object detection baselines on three datasets to show that DeGPR is model-agnostic, and consistently improves baselines obtaining up to 9% (absolute) mAP gains. Aayush Kumar Tyagi, Chirag Mohapatra, Prasenjit Das 0006, Govind Makharia, Lalita Mehra, Prathosh A. P., Mausam |
CVPR | 7 |
| 2023 | Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMsabstractA popular approach for improving the correctness of output from large language models (LLMs) is Self-Consistency -poll the LLM multiple times and output the most frequent solution.Existing Self-Consistency techniques always generate a constant number of samples per question, where a better approach will be to non-uniformly distribute the available budget based on the amount of agreement in the samples generated so far.In response, we introduce Adaptive-Consistency, a cost-efficient, model-agnostic technique that dynamically adjusts the number of samples per question using a lightweight stopping criterion.Our experiments over 17 reasoning and code generation datasets and three LLMs demonstrate that Adaptive-Consistency reduces sample budget by up to 7.9 times with an average accuracy drop of less than 0.1%. 1 Pranjal Aggarwal, Aman Madaan, Yiming Yang 0002, Mausam |
EMNLP | 4 |
| 2023 | Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language ModelsabstractThe performance of large language models (LLMs) on existing reasoning benchmarks has significantly improved over the past years.In response, we present JEEBENCH, a considerably more challenging benchmark dataset for evaluating the problem solving abilities of LLMs.We curate 515 challenging preengineering mathematics, physics and chemistry problems from the highly competitive IIT JEE-Advanced exam.Long-horizon reasoning on top of deep in-domain knowledge is essential for solving problems in this benchmark.Our evaluation on various open-source and proprietary models reveals that the highest performance, even after using techniques like self-consistency, self-refinement and chain-ofthought prompting, is less than 40%.The typical failure modes of GPT-4, the best model, are errors in algebraic manipulation, difficulty in grounding abstract concepts into mathematical equations accurately and failure in retrieving relevant domain-specific concepts.We also observe that by mere prompting, GPT-4 is unable to assess risk introduced by negative marking for incorrect answers.For this, we develop a post-hoc confidence-thresholding method over self-consistency, which enables effective response selection.We hope that our challenging benchmark will guide future re-search in problem-solving using LLMs. Daman Arora, Himanshu Gaurav Singh, Mausam |
EMNLP | 3 |
| 2023 | ZGUL: Zero-shot Generalization to Unseen Languages using Multi-source Ensembling of Language AdaptersabstractWe tackle the problem of zero-shot crosslingual transfer in NLP tasks via the use of language adapters (LAs).Most of the earlier works have explored training with adapter of a single source (often English), and testing either using the target LA or LA of another related language.Training target LA requires unlabeled data, which may not be readily available for low resource unseen languages: those that are neither seen by the underlying multilingual language model (e.g., mBERT), nor do we have any (labeled or unlabeled) data for them.We posit that for more effective cross-lingual transfer, instead of just one source LA, we need to leverage LAs of multiple (linguistically or geographically related) source languages, both at train and test-time -which we investigate via our novel neural architecture, ZGUL.Extensive experimentation across four language groups, covering 15 unseen target languages, demonstrates improvements of up to 3.2 average F1 points over standard fine-tuning and other strong baselines on POS tagging and NER tasks.We also extend ZGUL to settings where either (1) some unlabeled data or (2) few-shot training examples are available for the target language.We find that ZGUL continues to outperform baselines in these settings too. Vipul Rathore, Rajdeep Dhingra, Parag Singla, Mausam |
EMNLP | 4 |
| 2023 | NeuSTIP: A Neuro-Symbolic Model for Link and Time Prediction in Temporal Knowledge GraphsabstractNeuro-symbolic (NS) models for knowledge graph completion (KGC) combine the benefits of symbolic models (interpretable inference) with those of distributed representations (parameter sharing, high accuracy).While several NS models exist for KGs with static facts, there is limited work on temporal KGC (TKGC) for KGs where a fact is associated with a time interval.In response, we propose a novel NS model for TKGC called NeuSTIP, which performs link prediction and time interval prediction in a TKG.NeuSTIP learns temporal rules with Allen predicates, which ensure temporal consistency between neighboring predicates in the rule body.We further design a unique scoring function that evaluates the confidence of the candidate answers while performing link and time interval predictions by utilizing the learned rules.Our empirical evaluation on two time interval based TKGC datasets shows that our model shows competitive performance on link prediction and establishes a new state of the art on time prediction. Ishaan Singh, Garima Gaur, Mausam |
EMNLP | 4 |
| 2023 | SymNet 3.0: Exploiting Long-Range Influences in Learning Generalized Neural Policies for Relational MDPsabstractWe focus on the learning of generalized neural policies for Relational Markov Decision Processes (RMDPs) expressed in RDDL. Recent work first converts the instances of a relational domain into an instance graph, and then trains a Graph Attention Network (GAT) of fixed depth with parameters shared across instances to learn a state representation, which can be decoded to get the policy [sharma et al., 22]. Unfortunately, this approach struggles to learn policies that exploit long-range dependencies – a fact we formally prove in this paper. As a remedy, we first construct a novel influence graph characterized by edges capturing one-step influence (dependence) between nodes based on the transition model. We then define influence distance between two nodes as the shortest path between them in this graph – a feature we exploit to represent long-range dependencies. We show that our architecture, referred to as Symbolic Influence Network (SymNet3.0), with its distance-based features, does not suffer from the representational issues faced by earlier approaches. Extensive experimentation demonstrates that we are competitive with existing baselines on 12 standard IPPC domains, and perform significantly better on six additional domains (including IPPC variants), designed to test a model’s capability in capturing long-range dependencies. Further analysis shows that SymNet3.0 automatically learns to focus on nodes that have key information for representing policies that capture long-range dependencies. Daman Arora, Mausam, Parag Singla |
UAI | 3 |
| 2022 | Alignment-Augmented Consistent Translation for Multilingual Open Information ExtractionabstractProgress with supervised Open Information Extraction (OpenIE) has been primarily limited to English due to the scarcity of training data in other languages.In this paper, we explore techniques to automatically convert English text for training OpenIE systems in other languages.We introduce the Alignment-Augmented Consistent Translation (AACTRANS) model to translate English sentences and their corresponding extractions consistently with each other -with no changes to vocabulary or semantic meaning which may result from independent translations.Using the data generated with AACTRANS, we train a novel two-stage generative OpenIE model, which we call GEN2OIE, that outputs for each sentence: 1) relations in the first stage and 2) all extractions containing the relation in the second stage.GEN2OIE increases relation coverage using a training data transformation technique that is generalizable to multiple languages, in contrast to existing models that use an English-specific training loss.Evaluations on 5 languages -Spanish, Portuguese, Chinese, Hindi and Telugu -show that the GEN2OIE with AACTRANS data outperforms prior systems by a margin of 6-25% F1. 1 Keshav Kolluru, Muqeeth Mohammed, Shubham Mittal 0001, Soumen Chakrabarti, Mausam |
ACL (1) | 5 |
| 2022 | Joint Completion and Alignment of Multilingual Knowledge GraphsabstractKnowledge Graph Completion (KGC) predicts missing facts in an incomplete Knowledge Graph (KG).Multilingual KGs associate entities and relations with surface forms written in different languages.An entity or relation may be associated with distinct IDs in different KGs, necessitating entity alignment (EA) and relation alignment (RA).Many effective algorithms have been proposed for completion and alignment as separate tasks.Here we show that these tasks are synergistic and best solved together.Our multitask approach starts with a state-of-the-art KG embedding scheme, but adds a novel relation representation based on sets of embeddings of (subject, object) entity pairs.This representation leads to a new relation alignment loss term based on a maximal bipartite matching between two sets of embedding vectors.This loss is combined with traditional KGC loss and optionally, losses based on text embeddings of entity (and relation) names.In experiments over KGs in seven languages, we find that our system achieves large improvements in KGC compared to a strong completion model that combines known facts in all languages.It also outperforms strong EA and RA baselines, underscoring the value of joint alignment and completion. Soumen Chakrabarti, Harkanwar Singh, Shubham Lohiya, Prachi Jain 0001, Mausam |
EMNLP | 5 |
| 2022 | "Covid vaccine is against Covid but Oxford vaccine is made at Oxford!" Semantic Interpretation of Proper Noun CompoundsabstractProper noun compounds, e.g., "Covid vaccine", convey information in a succinct manner (a "Covid vaccine" is a "vaccine that immunizes against the Covid disease").These are commonly used in short-form domains, such as news headlines, but are largely ignored in information-seeking applications.To address this limitation, we release a new manually annotated dataset, PRONCI, consisting of 22.5K proper noun compounds along with their freeform semantic interpretations.PRONCI is 60 times larger than prior noun compound datasets and also includes non-compositional examples, which have not been previously explored.We experiment with various neural models for automatically generating the semantic interpretations from proper noun compounds, ranging from few-shot prompting to supervised learning, with varying degrees of knowledge about the constituent nouns.We find that adding targeted knowledge, particularly about the common noun, results in performance gains of upto 2.8%.Finally, we integrate our model generated interpretations with an existing Open IE system and observe an 7.5% increase in yield at a precision of 85%.The dataset and code are available at https://github.com/dair-iitd/pronci. Keshav Kolluru, Gabriel Stanovsky, Mausam |
EMNLP | 3 |
| 2022 | Structural Constraints and Natural Language Inference for End-to-End Flowchart Grounded Dialog Response GenerationabstractFlowchart grounded dialog systems converse with users by following a given flowchart and a corpus of FAQs.The existing state-of-the-art approach (Raghu et al., 2021) for learning such a dialog system, named FLONET, has two main limitations.(1) It uses a Retrieval Augmented Generation (RAG) framework which represents a flowchart as a bag of nodes.By doing so, it loses the connectivity structure between nodes which can aid in better response generation.(2) Typically dialogs progress with the agent asking polar (Y/N) questions, but users often respond indirectly without the explicit use of polar words.In such cases, it fails to understand the correct polarity of the answer.To overcome these issues, we propose Structure-Aware FLONET (SA-FLONET) which infuses structural constraints derived from the connectivity structure of flowcharts into the RAG framework.It uses natural language inference to better predict the polarity of indirect Y/N answers.We find that SA-FLONET outperforms FLONET, with a success rate improvement of 68% and 123% in flowchart grounded response generation and zero-shot flowchart grounded response generation tasks respectively. Dinesh Raghu, Suraj Joshi, Sachindra Joshi, Mausam |
EMNLP | 4 |
| 2022 | Neural Models for Output-Space Invariance in Combinatorial Problems
Yatin Nandwani, Vidit Jain, Mausam, Parag Singla |
ICLR | 3 |
| 2022 | A Solver-free Framework for Scalable Learning in Neural ILP ArchitecturesabstractThere is a recent focus on designing architectures that have an Integer Linear Programming (ILP) layer within a neural model (referred to as \emph{Neural ILP} in this paper). Neural ILP architectures are suitable for pure reasoning tasks that require data-driven constraint learning or for tasks requiring both perception (neural) and reasoning (ILP). A recent SOTA approach for end-to-end training of Neural ILP explicitly defines gradients through the ILP black box [Paulus et al. [2021]] – this trains extremely slowly, owing to a call to the underlying ILP solver for every training data point in a minibatch. In response, we present an alternative training strategy that is \emph{solver-free}, i.e., does not call the ILP solver at all at training time. Neural ILP has a set of trainable hyperplanes (for cost and constraints in ILP), together representing a polyhedron. Our key idea is that the training loss should impose that the final polyhedron separates the positives (all constraints satisfied) from the negatives (at least one violated constraint or a suboptimal cost value), via a soft-margin formulation. While positive example(s) are provided as part of the training data, we devise novel techniques for generating negative samples. Our solution is flexible enough to handle equality as well as inequality constraints. Experiments on several problems, both perceptual as well as symbolic, which require learning the constraints of an ILP, show that our approach has superior performance and scales much better compared to purely neural baselines and other state-of-the-art models that require solver-based training. In particular, we are able to obtain excellent performance in 9 x 9 symbolic and visual Sudoku, to which the other Neural ILP solver is not able to scale. Yatin Nandwani, Rishabh Ranjan, Mausam, Parag Singla |
NeurIPS | 3 |
| 2022 | SymNet 2.0: Effectively handling Non-Fluents and Actions in Generalized Neural Policies for RDDL Relational MDPsabstractRelational MDPs (RMDPs) compactly represent an infinite set of MDPs with an unbounded number of objects. Solving an RMDP requires a generalized policy that applies to all instances of a domain. Recently, Garg et al. proposed SymNet for this task– it constructs a graph neural network that shares parameters across all instances in a domain, thus making it applicable to any instance in a zero-shot manner. Our analysis of SymNet reveals that it performs no better than random on 1/4th of planning competition domains. The key reasons are its design choices: it misses important information during graph construction, leading to (1) poor generalizability, and (2) potential non-identifiability of different actions. In response, our solution, SymNet2.0, substantially augments SymNet’s graph construction approach by introducing additional nodes and edges which allow a better transfer of important information about a domain. It also improves SymNet’s action decoders with relevant information from objects to make different actions identifiable during scoring. Extensive experiments on twelve competition domains, where we use imitation learning over data generated from the PROST planner, demonstrate that SymNet2.0 performs vastly better than SymNet. Interestingly, even though SymNet2.0 is trained over data from PROST, it outperforms the planner on several test instances due to former’s ability to scale to large instances in a zero-shot manner. Daman Arora, Florian Geißer, Mausam, Parag Singla |
UAI | 4 |
| 2022 | TOOLTANGO: Common sense Generalization in Predicting Sequential Tool Interactions for Robot Plan SynthesisabstractRobots assisting us in environments such as factories or homes must learn to make use of objects as tools to perform tasks, for instance, using a tray to carry objects. We consider the problem of learning common sense knowledge of when a tool may be useful and how its use may be composed with other tools to accomplish a high-level task instructed by a human. Specifically, we introduce a novel neural model, termed TOOLTANGO, that first predicts the next tool to be used, and then uses this information to predict the next action. We show that this joint model can inform learning of a fine-grained policy enabling the robot to use a particular tool in sequence and adds a significant value in making the model more accurate. TOOLTANGO encodes the world state, comprising objects and symbolic relationships between them, using a graph neural network and is trained using demonstrations from human teachers instructing a virtual robot in a physics simulator. The model learns to attend over the scene using knowledge of the goal and the action history, finally decoding the symbolic action to execute. Crucially, we address generalization to unseen environments where some known tools are missing, but unseen alternative tools are present. We show that by augmenting the representation of the environment with pre-trained embeddings derived from a knowledge-base, the model can generalize effectively to novel environments. Experimental results show at least 48.8-58.1% absolute improvement over the baselines in predicting successful symbolic plans for a simulated mobile manipulator in novel environments with unseen objects. This work takes a step in the direction of enabling robots to rapidly synthesize robust plans for complex tasks, particularly in novel settings. Shreshth Tuli, Rajas Bansal, Rohan Paul, Mausam |
J. Artif. Intell. Res. | 4 |
| 2021 | Answering POI-recommendation Questions using Tourism ReviewsabstractWe introduce the novel and challenging task of answering Points-of-interest (POI) recommendation questions, using a collection of reviews that describe candidate answer entities (POIs). We harvest a QA dataset that contains 47,124 paragraph-sized user questions from travelers seeking POI recommendations for hotels, attractions and restaurants. Each question can have thousands of candidate entities to choose from and each candidate is associated with a collection of unstructured reviews. Questions can include requirements based on physical location, budget, timings as well as other subjective considerations related to ambience, quality of service etc. Our dataset requires reasoning over a large number of candidate answer entities (over 5300 per question on average) and we find that running commonly used neural architectures for QA is prohibitively expensive. Further, commonly used retriever-ranker based methods also do not work well for our task due to the nature of review-documents. Thus, as a first attempt at addressing some of the novel challenges of reasoning-at-scale posed by our task, we present a task specific baseline model that uses a three-stage cluster-select-rerank architecture. The model first clusters text for each entity to identify exemplar sentences describing an entity. It then uses a neural information retrieval (IR) module to select a set of potential entities from the large candidate set. A reranker uses a deeper attention-based architecture to pick the best answers from the selected entities. This strategy performs better than a pure retrieval or a pure attention-based reasoning approach yielding nearly 25% relative improvement in [email protected] over both approaches. To the best of our knowledge we are the first to present an unstructured QA-style task for POI-recommendation, using real-world tourism questions and POI-reviews. Danish Contractor, Krunal Shah 0001, Aditi Partap, Parag Singla, Mausam |
CIKM | 5 |
| 2021 | End-to-End Learning of Flowchart Grounded Task-Oriented DialogsabstractWe propose a novel problem within end-toend learning of task oriented dialogs (TOD), in which the dialog system mimics a troubleshooting agent who helps a user by diagnosing their problem (e.g., car not starting).Such dialogs are grounded in domain-specific flowcharts, which the agent is supposed to follow during the conversation.Our task exposes novel technical challenges for neural TOD, such as grounding an utterance to the flowchart without explicit annotation, referring to additional manual pages when user asks a clarification question, and ability to follow unseen flowcharts at test time.We release a dataset (FLODIAL) consisting of 2,738 dialogs grounded on 12 different troubleshooting flowcharts.We also design a neural model, FLONET, which uses a retrieval-augmented generation architecture to train the dialog agent.Our experiments find that FLONET can do zero-shot transfer to unseen flowcharts, and sets a strong baseline for future research. Dinesh Raghu, Shantanu Agarwal, Sachindra Joshi, Mausam |
EMNLP (1) | 4 |
| 2021 | Neural Learning of One-of-Many Solutions for Combinatorial Problems in Structured Output Spaces
Yatin Nandwani, Deepanshu Jindal, Mausam, Parag Singla |
ICLR | 3 |
| 2021 | TANGO: Commonsense Generalization in Predicting Tool Interactions for Mobile ManipulatorsabstractRobots assisting us in factories or homes must learn to make use of objects as tools to perform tasks, e.g., a tray for carrying objects. We consider the problem of learning commonsense knowledge of when a tool may be useful and how its use may be composed with other tools to accomplish a high-level task instructed by a human. We introduce TANGO, a novel neural model for predicting task-specific tool interactions. TANGO is trained using demonstrations obtained from human teachers instructing a virtual robot in a physics simulator. TANGO encodes the world state consisting of objects and symbolic relationships between them using a graph neural network. The model learns to attend over the scene using knowledge of the goal and the action history, finally decoding the symbolic action to execute. Crucially, we address generalization to unseen environments where some known tools are missing, but alternative unseen tools are present. We show that by augmenting the representation of the environment with pre-trained embeddings derived from a knowledge-base, the model can generalize effectively to novel environments. Experimental results show a 60.5-78.9% improvement over the baseline in predicting successful symbolic plans in unseen settings for a simulated mobile manipulator. Shreshth Tuli, Rajas Bansal, Rohan Paul, Mausam |
IJCAI | 4 |
| 2021 | Joint Spatio-Textual Reasoning for Answering Tourism QuestionsabstractOur goal is to answer real-world tourism questions that seek Points-of-Interest (POI) recommendations. Such questions express various kinds of spatial and non-spatial constraints, necessitating a combination of textual and spatial reasoning. In response, we develop the first joint spatio-textual reasoning model, which combines geo-spatial knowledge with information in textual corpora to answer questions. We first develop a modular spatial-reasoning network that uses geo-coordinates of location names mentioned in a question, and of candidate answer POIs, to reason over only spatial constraints. We then combine our spatial-reasoner with a textual reasoner in a joint model and present experiments on a real world POI recommendation task. We report substantial improvements over existing models without joint spatio-textual reasoning. To the best of our knowledge, we are the first to develop a joint QA model that combines reasoning over external geo-spatial knowledge along with textual reasoning. Danish Contractor, Shashank Goel, Mausam, Parag Singla |
WWW | 3 |
| 2021 | Constrained BERT BiLSTM CRF for understanding multi-sentence entity-seeking questionsabstractAbstract We present the novel task of understanding multi-sentenceentity-seekingquestions (MSEQs), that is, the questions that may be expressed in multiple sentences, and that expect one or more entities as an answer. We formulate the problem of understanding MSEQs as a semantic labeling task over an open representation that makes minimal assumptions about schema or ontology-specific semantic vocabulary. At the core of our model, we use a BiLSTM (bidirectional LSTM) conditional random field (CRF), and to overcome the challenges of operating with low training data, we supplement it by using BERT embeddings, hand-designed features, as well as hard and soft constraints spanning multiple sentences. We find that this results in a 12–15 points gain over a vanilla BiLSTM CRF. We demonstrate the strengths of our work using the novel task of answering real-world entity-seeking questions from the tourism domain. The use of our labels helps answer 36% more questions with 35% more (relative) accuracy as compared to baselines. We also demonstrate how our framework can rapidly enable the parsing of MSEQs in an entirely new domain with small amounts of training data and little change in the semantic representation. Danish Contractor, Barun Patra, Mausam, Parag Singla |
Nat. Lang. Eng. | 3 |
| 2021 | Unsupervised Learning of KB Queries in Task-Oriented DialogsabstractAbstract Task-oriented dialog (TOD) systems often need to formulate knowledge base (KB) queries corresponding to the user intent and use the query results to generate system responses. Existing approaches require dialog datasets to explicitly annotate these KB queries—these annotations can be time consuming, and expensive. In response, we define the novel problems of predicting the KB query and training the dialog agent, without explicit KB query annotation. For query prediction, we propose a reinforcement learning (RL) baseline, which rewards the generation of those queries whose KB results cover the entities mentioned in subsequent dialog. Further analysis reveals that correlation among query attributes in KB can significantly confuse memory augmented policy optimization (MAPO), an existing state of the art RL agent. To address this, we improve the MAPO baseline with simple but important modifications suited to our task. To train the full TOD system for our setting, we propose a pipelined approach: it independently predicts when to make a KB query (query position predictor), then predicts a KB query at the predicted position (query predictor), and uses the results of predicted query in subsequent dialog (next response predictor). Overall, our work proposes first solutions to our novel problem, and our analysis highlights the research challenges in training TOD systems without query annotation. Dinesh Raghu, Nikhil Gupta 0007, Mausam |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | IMoJIE: Iterative Memory-Based Joint Open Information ExtractionabstractWhile traditional systems for Open Information Extraction were statistical and rule-based, recently neural models have been introduced for the task.Our work builds upon CopyAttention, a sequence generation OpenIE model (Cui et al., 2018).Our analysis reveals that CopyAttention produces a constant number of extractions per sentence, and its extracted tuples often express redundant information.We present IMOJIE, an extension to Copy-Attention, which produces the next extraction conditioned on all previously extracted tuples.This approach overcomes both shortcomings of CopyAttention, resulting in a variable number of diverse extractions per sentence.We train IMOJIE on training data bootstrapped from extractions of several non-neural systems, which have been automatically filtered to reduce redundancy and noise.IMOJIE outperforms CopyAttention by about 18 F1 pts, and a BERT-based strong baseline by 2 F1 pts, establishing a new state of the art for the task. Keshav Kolluru, Samarth Aggarwal, Vipul Rathore, Mausam, Soumen Chakrabarti |
ACL | 4 |
| 2020 | A Simple Yet Strong Pipeline for HotpotQAabstractState-of-the-art models for multi-hop question answering typically augment large-scale language models like BERT with additional, intuitively useful capabilities such as named entity recognition, graph-based reasoning, and question decomposition.However, does their strong performance on popular multihop datasets really justify this added design complexity?Our results suggest that the answer may be no, because even our simple pipeline based on BERT, named QUARK, performs surprisingly well.Specifically, on Hot-potQA, QUARK outperforms these models on both question answering and support identification (and achieves performance very close to a RoBERTa model).Our pipeline has three steps: 1) use BERT to identify potentially relevant sentences independently of each other; 2) feed the set of selected sentences as context into a standard BERT span prediction model to choose an answer; and 3) use the sentence selection model, now with the chosen answer, to produce supporting sentences.The strong performance of QUARK resurfaces the importance of carefully exploring simple model designs before using popular benchmarks to justify the value of complex techniques. Dirk Groeneveld, Tushar Khot, Mausam, Ashish Sabharwal |
EMNLP (1) | 3 |
| 2020 | Temporal Knowledge Base Completion: New Algorithms and Evaluation ProtocolsabstractResearch on temporal knowledge bases, which associate a relational fact (s, r, o) with a validity time period (or time instant), is in its early days.Our work considers predicting missing entities (link prediction) and missing time intervals (time prediction) as joint Temporal Knowledge Base Completion (TKBC) tasks, and presents TIMEPLEX, a novel TKBC method, in which entities, relations and, time are all embedded in a uniform, compatible space.TIMEPLEX exploits the recurrent nature of some facts/events and temporal interactions between pairs of relations, yielding stateof-the-art results on both prediction tasks.We also find that existing TKBC models heavily overestimate link prediction performance due to imperfect evaluation mechanisms.In response, we propose improved TKBC evaluation protocols for both link and time prediction tasks, dealing with subtle issues that arise from the partial overlap of time intervals in gold instances and system predictions. Prachi Jain 0001, Sushant Rathi, Mausam, Soumen Chakrabarti |
EMNLP (1) | 3 |
| 2020 | OpenIE6: Iterative Grid Labeling and Coordination Analysis for Open Information ExtractionabstractA recent state-of-the-art neural open information extraction (OpenIE) system generates extractions iteratively, requiring repeated encoding of partial outputs.This comes at a significant computational cost.On the other hand, sequence labeling approaches for OpenIE are much faster, but worse in extraction quality.In this paper, we bridge this trade-off by presenting an iterative labeling-based system that establishes a new state of the art for OpenIE, while extracting 10× faster.This is achieved through a novel Iterative Grid Labeling (IGL) architecture, which treats OpenIE as a 2-D grid labeling task.We improve its performance further by applying coverage (soft) constraints on the grid at training time.Moreover, on observing that the best OpenIE systems falter at handling coordination structures, our OpenIE system also incorporates a new coordination analyzer built with the same IGL architecture.This IGL based coordination analyzer helps our OpenIE system handle complicated coordination structures, while also establishing a new state of the art on the task of coordination analysis, with a 12.3 pts improvement in F1 over previous analyzers.Our OpenIE system, OpenIE6 1 , beats the previous systems by as much as 4 pts in F1, while being much faster. Keshav Kolluru, Vaibhav Adlakha, Samarth Aggarwal, Mausam, Soumen Chakrabarti |
EMNLP (1) | 4 |
| 2020 | Symbolic Network: Generalized Neural Policies for Relational MDPsabstractA Relational Markov Decision Process (RMDP) is a first-order representation to express all instances of a single probabilistic planning domain with possibly unbounded number of objects. Early work in RMDPs outputs generalized (instance-independent) first-order policies or value functions as a means to solve all instances of a domain at once. Unfortunately, this line of work met with limited success due to inherent limitations of the representation space used in such policies or value functions. Can neural models provide the missing link by easily representing more complex generalized policies, thus making them effective on all instances of a given domain? We present SymNet, the first neural approach for solving RMDPs that are expressed in the probabilistic planning language of RDDL. SymNet trains a set of shared parameters for an RDDL domain using training instances from that domain. For each instance, SymNet first converts it to an instance graph and then uses relational neural models to compute node embeddings. It then scores each ground action as a function over the first-order action symbols and node embeddings related to the action. Given a new test instance from the same domain, SymNet architecture with pre-trained parameters scores each ground action and chooses the best action. This can be accomplished in a single forward pass without any retraining on the test instance, thus implicitly representing a neural generalized policy for the whole domain. Our experiments on nine RDDL domains from IPPC demonstrate that SymNet policies are significantly better than random and sometimes even more effective than training a state-of-the-art deep reactive policy from scratch. Sankalp Garg, Aniket Bajpai, Mausam |
ICML | 3 |
| 2019 | CaRB: A Crowdsourced Benchmark for Open IEabstractSangnie Bhardwaj, Samarth Aggarwal, Mausam Mausam. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sangnie Bhardwaj, Samarth Aggarwal, Mausam |
EMNLP/IJCNLP (1) | 3 |
| 2019 | A Primal Dual Formulation For Deep Learning With ConstraintsabstractFor several problems of interest, there are natural constraints which exist over the output label space. For example, for the joint task of NER and POS labeling, these constraints might specify that the NER label ‘organization’ is consistent only with the POS labels ‘noun’ and ‘preposition’. These constraints can be a great way of injecting prior knowledge into a deep learning model, thereby improving overall performance. In this paper, we present a constrained optimization formulation for training a deep network with a given set of hard constraints on output labels. Our novel approach first converts the label constraints into soft logic constraints over probability distributions outputted by the network. It then converts the constrained optimization problem into an alternating min-max optimization with Lagrangian variables defined for each constraint. Since the constraints are independent of the target labels, our framework easily generalizes to semi-supervised setting. We experiment on the tasks of Semantic Role Labeling (SRL), Named Entity Recognition (NER) tagging, and fine-grained entity typing and show that our constraints not only significantly reduce the number of constraint violations, but can also result in state-of-the-art performance Yatin Nandwani, Abhishek Pathak, Mausam, Parag Singla |
NeurIPS | 3 |
| 2018 | Open Information Extraction from Conjunctive SentencesabstractWe develop CALM, a coordination analyzer that improves upon the conjuncts identified from dependency parses. It uses a language model based scoring and several linguistic constraints to search over hierarchical conjunct boundaries (for nested coordination). By splitting a conjunctive sentence around these conjuncts, CALM outputs several simple sentences. We demonstrate the value of our coordination analyzer in the end task of Open Information Extraction (Open IE). State-of-the-art Open IE systems lose substantial yield due to ineffective processing of conjunctive sentences. Our Open IE system, CALMIE, performs extraction over the simple sentences identified by CALM to obtain up to 1.8x yield with a moderate increase in precision compared to extractions from original sentences. Swarnadeep Saha, Mausam |
COLING | 2 |
| 2018 | Active Learning with Unbalanced Classes and Example-Generation QueriesabstractMachine learning in real-world high-skew domains is difficult, because traditional strategies for crowdsourcing labeled training examples are ineffective at locating the scarce minority-class examples. For example, both random sampling and traditional active learning (which reduces to random sampling when just starting) will most likely recover very few minority-class examples. To bootstrap the machine learning process, researchers have proposed tasking the crowd with finding or generating minority-class examples, but such strategies have their weaknesses as well. They are unnecessarily expensive in well-balanced domains, and they often yield samples from a biased distribution that is unrepresentative of the one being learned.This paper extends the traditional active learning framework by investigating the problem of intelligently switching between various crowdsourcing strategies for obtaining labeled training examples in order to optimally train a classifier. We start by analyzing several such strategies (e.g., annotate an example, generate a minority-class example, etc.), and then develop a novel, skew-robust algorithm, called MB-CB, for the control problem. Experiments show that our method outperforms state-of-the-art GL-Hybrid by up to 14.3 points in F1 AUC, across various domains and class-frequency settings. Christopher H. Lin, Mausam, Daniel S. Weld |
HCOMP | 2 |
| 2018 | Mitigating the Effect of Out-of-Vocabulary Entity Pairs in Matrix Factorization for KB InferenceabstractThis paper analyzes the varied performance of Matrix Factorization (MF) on the related tasks of relation extraction and knowledge-base completion, which have been unified recently into a single framework of knowledge-base inference (KBI) [Toutanova et al., 2015]. We first propose a new evaluation protocol that makes comparisons between MF and Tensor Factorization (TF) models fair. We find that this results in a steep drop in MF performance. Our analysis attributes this to the high out-of-vocabulary (OOV) rate of entity pairs in test folds of commonly-used datasets. To alleviate this issue, we propose three extensions to MF. Our best model is a TF-augmented MF model. This hybrid model is robust and obtains strong results across various KBI datasets. Prachi Jain 0001, Shikhar Murty, Mausam, Soumen Chakrabarti |
IJCAI | 3 |
| 2018 | Inferring Temporal Knowledge for Near-Periodic Recurrent EventsabstractWe define the novel problem of extracting and predicting occurrence dates for a class of recurrent events -- events that are held periodically as per a near-regular schedule (e.g., conferences, film festivals, sport championships). Knowledge-bases such as Freebase contain a large number of such recurring events, but they also miss substantial information regarding specific event instances and their occurrence dates. We develop a temporal extraction and inference engine to fill in the missing dates as well as to predict their future occurrences. Our engine performs joint inference over several knowledge sources -- (1) information about an event instance and its date extracted from text by our temporal extractor, (2) information about the typical schedule (e.g., ``every second week of June") for a recurrent event extracted by our schedule extractor, and (3) known dates for other instances of the same event. The output of our system is a representation for the event schedule and an occurrence date for each event instance. We find that our system beats humans in predicting future occurrences of recurrent events by significant margins. We release our code and system output for further research. Dinesh Raghu, Surag Nair, Mausam |
IJCAI | 3 |
| 2018 | Transfer of Deep Reactive Policies for MDP PlanningabstractDomain-independent probabilistic planners input an MDP description in a factored representation language such as PPDDL or RDDL, and exploit the specifics of the representation for faster planning. Traditional algorithms operate on each problem instance independently, and good methods for transferring experience from policies of other instances of a domain to a new instance do not exist. Recently, researchers have begun exploring the use of deep reactive policies, trained via deep reinforcement learning (RL), for MDP planning domains. One advantage of deep reactive policies is that they are more amenable to transfer learning. In this paper, we present the first domain-independent transfer algorithm for MDP planning domains expressed in an RDDL representation. Our architecture exploits the symbolic state configuration and transition function of the domain (available via RDDL) to learn a shared embedding space for states and state-action pairs for all problem instances of a domain. We then learn an RL agent in the embedding space, making a near zero-shot transfer possible, i.e., without much training on the new instance, and without using the domain simulator at all. Experiments on three different benchmark domains underscore the value of our transfer algorithm. Compared against planning from scratch, and a state-of-the-art RL transfer algorithm, our transfer solution has significantly superior learning curves. Aniket Bajpai, Sankalp Garg, Mausam |
NeurIPS | 3 |
| 2018 | Block-Value Symmetries in Probabilistic Graphical Models
Gagan Madan, Ankit Anand, Mausam, Parag Singla |
UAI | 3 |
| 2018 | Sprout: Crowd-Powered Task Design for CrowdsourcingabstractWhile crowdsourcing enables data collection at scale, ensuring high-quality data remains a challenge. In particular, effective task design underlies nearly every reported crowdsourcing success, yet remains difficult to accomplish. Task design is hard because it involves a costly iterative process: identifying the kind of work output one wants, conveying this information to workers, observing worker performance, understanding what remains ambiguous, revising the instructions, and repeating the process until the resulting output is satisfactory. To facilitate this process, we propose a novel meta-workflow that helps requesters optimize crowdsourcing task designs and Sprout, our open-source tool, which implements this workflow. Sprout improves task designs by (1) eliciting points of confusion from crowd workers, (2) enabling requesters to quickly understand these misconceptions and the overall space of questions, and (3) guiding requesters to improve the task design in response. We report the results of a user study with two labeling tasks demonstrating that requesters strongly prefer Sprout and produce higher-rated instructions compared to current best practices for creating gated instructions (instructions plus a workflow for training and testing workers). We also offer a set of design recommendations for future tools that support crowdsourcing task design. Jonathan Bragg, Mausam, Daniel S. Weld |
UIST | 2 |
| 2017 | Non-Count Symmetries in Boolean & Multi-Valued Prob. Graphical ModelsabstractLifted inference algorithms commonly exploit symmetries in a probabilistic graphical model (PGM) for efficient inference. However, existing algorithms for Boolean-valued domains can identify only those pairs of states as symmetric, in which the number of ones and zeros match exactly (count symmetries). Moreover, algorithms for lifted inference in multi-valued domains also compute a multi-valued extension of count symmetries only. These algorithms miss many symmetries in a domain. In this paper, we present first algorithms to compute non-count symmetries in both Boolean-valued and multi-valued domains. Our methods can also find symmetries between multi-valued variables that have different domain cardinalities. The key insight in the algorithms is that they change the unit of symmetry computation from a variable to a variable-value (VV) pair. Our experiments find that exploiting these symmetries in MCMC can obtain substantial computational gains over existing algorithms. Ankit Anand, Ritesh Noothigattu, Parag Singla, Mausam |
AISTATS | 4 |
| 2017 | Octopus: A Framework for Cost-Quality-Time Optimization in CrowdsourcingabstractWe present Octopus, an AI agent to jointly balance three conflicting task objectives on a micro-crowdsourcing marketplace – the quality of work, total cost incurred, and time to completion. Previous control agents have mostly focused on cost-quality, or cost-time tradeoffs, but not on directly controlling all three in concert. A naive formulation of three-objective optimization is intractable; Octopus takes a hierarchical POMDP approach, with three different components responsible for setting the pay per task, selecting the next task, and controlling task-level quality. We demonstrate that Octopus significantly outperforms existing state-of-the-art approaches on real experiments. We also deploy Octopus on Amazon Mechanical Turk, showing its ability to manage tasks in a real-world, dynamic setting. Karan Goel, Shreya Rajpal, Mausam |
HCOMP | 3 |
| 2017 | Coarse-to-Fine Lifted MAP Inference in Computer VisionabstractThere is a vast body of theoretical research on lifted inference in probabilistic graphical models (PGMs). However, few demonstrations exist where lifting is applied in conjunction with top of the line applied algorithms. We pursue the applicability of lifted inference for computer vision (CV), with the insight that a globally optimal (MAP) labeling will likely have the same label for two symmetric pixels. The success of our approach lies in efficiently handling a distinct unary potential on every node (pixel), typical of CV applications. This allows us to lift the large class of algorithms that model a CV problem via PGM inference. We propose a generic template for coarse-to-fine (C2F) inference in CV, which progressively refines an initial coarsely lifted PGM for varying quality-time trade-offs. We demonstrate the performance of C2F inference by developing lifted versions of two near state-of-the-art CV algorithms for stereo vision and interactive image segmentation. We find that, against flat algorithms, the lifted versions have a much superior anytime performance, without any loss in final solution quality. Haroun Habeeb, Ankit Anand, Mausam, Parag Singla |
IJCAI | 3 |
| 2016 | Re-Active Learning: Active Learning with RelabelingabstractActive learning seeks to train the best classifier at the lowest annotation cost by intelligently picking the best examples to label. Traditional algorithms assume there is a single annotator and disregard the possibility of requesting additional independent annotations for a previously labeled example. However, relabeling examples is important, because all annotators make mistakes — especially crowdsourced workers, who have become a common source of training data. This paper seeks to understand the difference in marginal value between decreasing the noise of the training set via relabeling and increasing the size and diversity of the (noisier) training set by labeling new examples. We use the term re-active learning to denote this generalization of active learning. We show how traditional active learning methods perform poorly at re-active learning, present new algorithms designed for this important problem, formally characterize their behavior, and empirically show that our methods effectively make this tradeoff. Christopher H. Lin, Mausam, Daniel S. Weld |
AAAI | 2 |
| 2016 | Numerical Relation Extraction with Minimal SupervisionabstractWe study a novel task of numerical relation extraction with the goal of extracting relations where one of the arguments is a number or a quantity ( e.g., atomic_number(Aluminium, 13), inflation_rate(India, 10.9%)). This task presents peculiar challenges not found in standard IE, such as the difficulty of matching numbers in distant supervision and the importance of units. We design two extraction systems that require minimal human supervision per relation: (1) NumberRule, a rule based extractor, and (2) NumberTron, a probabilistic graphical model. We find that both systems dramatically outperform MultiR, a state-of-the-art non-numerical IE model, obtaining up to 25 points F-score improvement. Aman Madaan, Ashish R. Mittal, Mausam, Ganesh Ramakrishnan, Sunita Sarawagi |
AAAI | 3 |
| 2016 | Contextual Symmetries in Probabilistic Graphical Models
Ankit Anand, Aditya Grover, Mausam, Parag Singla |
IJCAI | 3 |
| 2016 | Open Information Extraction Systems and Downstream Applications
Mausam |
IJCAI | 1 |
| 2016 | Entity-balanced Gaussian pLSA for Automated Comparison
Danish Contractor, Parag Singla, Mausam |
HLT-NAACL | 3 |
| 2016 | Knowledge-Guided Linguistic Rewrites for Inference Rule VerificationabstractA corpus of inference rules between a pair of relation phrases is typically generated using the statistical overlap of argument-pairs associated with the relations (e.g., PATTY, CLEAN).We investigate knowledge-guided linguistic rewrites as a secondary source of evidence and find that they can vastly improve the quality of inference rule corpora, obtaining 27 to 33 point precision improvement while retaining substantial recall.The facts inferred using cleaned inference rules are 29-32 points more accurate. Prachi Jain 0001, Mausam |
HLT-NAACL | 2 |
| 2015 | ASAP-UCT: Abstraction of State-Action Pairs in UCT
Ankit Anand, Aditya Grover, Mausam, Parag Singla |
IJCAI | 3 |
| 2014 | Hierarchical Summarization: Scaling Up Multi-Document SummarizationabstractMulti-document summarization (MDS) systems have been designed for short, un-structured summaries of 10-15 documents, and are inadequate for larger document collections. We propose a new approach to scaling up summarization called hierar-chical summarization, and present the first implemented system, SUMMA. SUMMA produces a hierarchy of relatively short summaries, in which the top level provides a general overview and users can navigate the hierarchy to drill down for more details on topics of interest. SUMMA optimizes for coherence as well as cover-age of salient information. In an Amazon Mechanical Turk evaluation, users pref-ered SUMMA ten times as often as flat MDS and three times as often as timelines. 1 Janara Christensen, Stephen Soderland, Gagan Bansal, Mausam |
ACL (1) | 4 |
| 2014 | Parallel Task Routing for CrowdsourcingabstractAn ideal crowdsourcing or citizen-science system would route tasks to the most appropriate workers, but the best assignment is unclear because workers have varying skill, tasks have varying difficulty, and assigning several workers to a single task may significantly improve output quality. This paper defines a space of task routing problems, proves that even the simplest is NP-hard, and develops several approximation algorithms for parallel routing problems. We show that an intuitive class of requesters' utility functions is submodular, which lets us provide iterative methods for dynamically allocating batches of tasks that make near-optimal use of available workers in each round. Experiments with live oDesk workers show that our task routing algorithm uses only 48% of the human labor compared to the commonly used round-robin strategy. Further, we provide versions of our task routing algorithm which enable it to scale to large numbers of workers and questions and to handle workers with variable response times while still providing significant benefit over common baselines. Jonathan Bragg, Andrey Kolobov, Mausam, Daniel S. Weld |
HCOMP | 3 |
| 2014 | To Re(label), or Not To Re(label)abstractOne of the most popular uses of crowdsourcing is to provide training data for supervised machine learning algorithms. Since human annotators often make errors, requesters commonly ask multiple workers to label each example. But is this strategy always the most cost effective use of crowdsourced workers? We argue "No" --- often classifiers can achieve higher accuracies when trained with noisy "unilabeled" data. However, in some cases relabeling is extremely important. We discuss three factors that may make relabeling an effective strategy: classifier expressiveness, worker accuracy, and budget. Christopher H. Lin, Mausam, Daniel S. Weld |
HCOMP | 2 |
| 2013 | Generating Coherent Event Schemas at ScaleabstractChambers and Jurafsky (2009) demonstrated that event schemas can be automatically induced from text corpora.However, our analysis of their schemas identifies several weaknesses, e.g., some schemas lack a common topic and distinct roles are incorrectly mixed into a single actor.It is due in part to their pair-wise representation that treats subjectverb independently from verb-object.This often leads to subject-verb-object triples that are not meaningful in the real-world.We present a novel approach to inducing open-domain event schemas that overcomes these limitations.Our approach uses cooccurrence statistics of semantically typed relational triples, which we call Rel-grams (relational n-grams).In a human evaluation, our schemas outperform Chambers's schemas by wide margins on several evaluation criteria.Both Rel-grams and event schemas are freely available to the research community. Niranjan Balasubramanian, Stephen Soderland, Mausam, Oren Etzioni |
EMNLP | 3 |
| 2013 | Crowdsourcing Multi-Label Classification for Taxonomy CreationabstractRecent work has introduced CASCADE, an algorithm for creating a globally-consistent taxonomy by crowdsourcing microwork from many individuals, each of whom may see only a tiny fraction of the data (Chilton et al. 2013). While CASCADE needs only unskilled labor and produces taxonomies whose quality approaches that of human experts, it uses significantly more labor than experts. This paper presents DELUGE, an improved workflow that produces taxonomies with comparable quality using significantly less crowd labor. Specifically, our method for crowdsourcing multi-label classification optimizes CASCADE’s most costly step (categorization) using less than 10% of the labor required by the original approach. DELUGE’s savings come from the use of decision theory and machine learning, which allow it to pose microtasks that aim to maximize information gain. Jonathan Bragg, Mausam, Daniel S. Weld |
HCOMP | 2 |
| 2013 | Towards Coherent Multi-Document Summarization
Janara Christensen, Mausam, Stephen Soderland, Oren Etzioni |
HLT-NAACL | 2 |
| 2013 | POMDP-based control of workflows for crowdsourcing
Peng Dai 0001, Christopher H. Lin, Mausam, Daniel S. Weld |
Artif. Intell. | 3 |
| 2013 | Modeling Missing Data in Distant Supervision for Information ExtractionabstractDistant supervision algorithms learn information extraction models given only large readily available databases and text collections. Most previous work has used heuristics for generating labeled data, for example assuming that facts not contained in the database are not mentioned in the text, and facts in the database must be mentioned at least once. In this paper, we propose a new latent-variable approach that models missing data. This provides a natural way to incorporate side information, for instance modeling the intuition that text will often mention rare entities which are likely to be missing in the database. Despite the added complexity introduced by reasoning about missing data, we demonstrate that a carefully designed local search approach to inference is very accurate and scales to large datasets. Experiments demonstrate improved performance for binary and unary relation extraction when compared to learning with heuristic labels, including on average a 27% increase in area under the precision recall curve in the binary case. Alan Ritter, Luke Zettlemoyer, Mausam, Oren Etzioni |
Trans. Assoc. Comput. Linguistics | 3 |
| 2012 | LRTDP Versus UCT for Online Probabilistic PlanningabstractUCT, the premier method for solving games such as Go, is also becoming the dominant algorithm for probabilistic planning. Out of the five solvers at the International Probabilistic Planning Competition (IPPC) 2011, four were based on the UCT algorithm. However, while a UCT-based planner, PROST, won the contest, an LRTDP-based system, Glutton, came in a close second, outperforming other systems derived from UCT. These results raise a question: what are the strengths and weaknesses of LRTDP and UCT in practice? This paper starts answering this question by contrasting the two approaches in the context of finite-horizon MDPs. We demonstrate that in such scenarios, UCT's lack of a sound termination condition is a serious practical disadvantage. In order to handle an MDP with a large finite horizon under a time constraint, UCT forces an expert to guess a non-myopic lookahead value for which it should be able to converge on the encountered states. Mistakes in setting this parameter can greatly hurt UCT's performance. In contrast, LRTDP's convergence criterion allows for an iterative deepening strategy. Using this strategy, LRTDP automatically finds the largest lookahead value feasible under the given time constraint. As a result, LRTDP has better performance and stronger theoretical properties. We present an online version of Glutton, named Gourmand, that illustrates this analysis and outperforms PROST on the set of IPPC-2011 problems. Andrey Kolobov, Mausam, Daniel S. Weld |
AAAI | 2 |
| 2012 | Dynamically Switching between Synergistic Workflows for CrowdsourcingabstractTo ensure quality results from unreliable crowdsourced workers, task designers often construct complex workflows and aggregate worker responses from redundant runs. Frequently, they experiment with several alternative workflows to accomplish the task, and eventually deploy the one that achieves the best performance during early trials. Surprisingly, this seemingly natural design paradigm does not achieve the full potential of crowdsourcing. In particular, using a single workflow (even the best) to accomplish a task is suboptimal. We show that alternative workflows can compose synergistically to yield much higher quality output. We formalize the insight with a novel probabilistic graphical model. Based on this model, we design and implement AGENTHUNT, a POMDP-based controller that dynamically switches between these workflows to achieve higher returns on investment. Additionally, we design offline and online methods for learning model parameters. Live experiments on Amazon Mechanical Turk demonstrate the superiority of AGENTHUNT for the task of generating NLP training data, yielding up to 50% error reduction and greater net utility compared to previous methods. Christopher H. Lin, Mausam, Daniel S. Weld |
AAAI | 2 |
| 2012 | No Noun Phrase Left Behind: Detecting and Typing Unlinkable Entities
Thomas Lin, Mausam, Oren Etzioni |
EMNLP-CoNLL | 2 |
| 2012 | Open Language Learning for Information Extraction
Mausam, Michael Schmitz 0002, Stephen Soderland, Robert Bart, Oren Etzioni |
EMNLP-CoNLL | 1 |
| 2012 | Open domain event extraction from twitterabstractTweets are the most up-to-date and inclusive stream of in- formation and commentary on current events, but they are also fragmented and noisy, motivating the need for systems that can extract, aggregate and categorize important events. Previous work on extracting structured representations of events has focused largely on newswire text; Twitter's unique characteristics present new challenges and opportunities for open-domain event extraction. This paper describes TwiCal-- the first open-domain event-extraction and categorization system for Twitter. We demonstrate that accurately extracting an open-domain calendar of significant events from Twitter is indeed feasible. In addition, we present a novel approach for discovering important event categories and classifying extracted events based on latent variable models. By leveraging large volumes of unlabeled data, our approach achieves a 14% increase in maximum F1 over a supervised baseline. A continuously updating demonstration of our system can be viewed at http://statuscalendar.com; Our NLP tools are available at http://github.com/aritter/ twitter_nlp. Alan Ritter, Mausam, Oren Etzioni, Sam Clark |
KDD | 2 |
| 2012 | A Theory of Goal-Oriented MDPs with Dead Ends
Andrey Kolobov, Mausam, Daniel S. Weld |
UAI | 2 |
| 2012 | Crowdsourcing Control: Moving Beyond Multiple Choice
Christopher H. Lin, Mausam, Daniel S. Weld |
UAI | 2 |
| 2012 | Discovering hidden structure in factored MDPs
Andrey Kolobov, Mausam, Daniel S. Weld |
Artif. Intell. | 2 |
| 2011 | Artificial Intelligence for Artificial Artificial IntelligenceabstractCrowdsourcing platforms such as Amazon Mechanical Turk have become popular for a wide variety of human intelligence tasks; however, quality control continues to be a significant challenge. Recently, we propose TurKontrol, a theoretical model based on POMDPs to optimize iterative, crowd-sourced workflows. However, they neither describe how to learn the model parameters, nor show its effectiveness in a real crowd-sourced setting. Learning is challenging due to the scale of the model and noisy data: there are hundreds of thousands of workers with high-variance abilities. This paper presents an end-to-end system that first learns TurKontrol's POMDP parameters from real Mechanical Turk data, and then applies the model to dynamically optimize live tasks. We validate the model and use it to control a successive-improvement process on Mechanical Turk. By modeling worker accuracy and voting patterns, our system produces significantly superior artifacts compared to those generated through nonadaptive workflows using the same amount of money. Peng Dai 0001, Mausam, Daniel S. Weld |
AAAI | 2 |
| 2011 | Named Entity Recognition in Tweets: An Experimental Study
Alan Ritter, Sam Clark, Mausam, Oren Etzioni |
EMNLP | 3 |
| 2011 | Open Information Extraction: The Second GenerationabstractHow do we scale information extraction to the massive size and unprecedented heterogeneity of the Web corpus? Beginning in 2003, our KnowItAll project has sought to extract high-quality knowledge from the Web. In 2007, we introduced the Open Information Extraction (Open IE) paradigm which eschews handlabeled training examples, and avoids domainspecific verbs and nouns, to develop unlexicalized, domain-independent extractors that scale to the Web corpus. Open IE systems have extracted billions of assertions as the basis for both commonsense knowledge and novel question-answering systems. This paper describes the second generation of Open IE systems, which rely on a novel model of how relations and their arguments are expressed in English sentences to double precision/recall compared with previous systems such as TEXTRUNNER and WOE. 1 Oren Etzioni, Anthony Fader, Janara Christensen, Stephen Soderland, Mausam |
IJCAI | 5 |
| 2011 | Towards Scalable MDP Algorithms
Andrey Kolobov, Mausam, Daniel S. Weld |
IJCAI | 2 |
| 2011 | An analysis of open information extraction based on semantic role labelingabstractOpen Information Extraction extracts relations from text without requiring a pre-specified domain or vocabulary. While existing techniques have used only shallow syntactic features, we investigate the use of semantic role labeling techniques for the task of Open IE. Semantic role labeling (SRL) and Open IE, although developed mostly in isolation, are quite related. We compare SRL-based open extractors, which perform computationally expensive, deep syntactic analysis, with TextRunner, an open extractor, which uses shallow syntactic analysis but is able to analyze many more sentences in a fixed amount of time and thus exploit corpus-level statistics. Our evaluation answers questions regarding these systems, including, can SRL extractors, which are trained on PropBank, cope with heterogeneous text found on the Web? Which extractor attains better precision, recall, f-measure, or running time? How does extractor performance vary for binary, n-ary and nested relations? How much do we gain by running multiple extractors? How do we select the optimal extractor given amount of data, available time, types of extractions desired? Janara Christensen, Mausam, Stephen Soderland, Oren Etzioni |
K-CAP | 2 |
| 2011 | Topological Value Iteration Algorithms
Peng Dai 0001, Mausam, Daniel S. Weld, Judy Goldsmith |
J. Artif. Intell. Res. | 2 |
| 2010 | Decision-Theoretic Control of Crowd-Sourced WorkflowsabstractCrowd-sourcing is a recent framework in which human intelligence tasks are outsourced to a crowd of unknown people ("workers") as an open call (e.g., on Amazon's Mechanical Turk). Crowd-sourcing has become immensely popular with hoards of employers ("requesters"), who use it to solve a wide variety of jobs, such as dictation transcription, content screening, etc. In order to achieve quality results, requesters often subdivide a large task into a chain of bite-sized subtasks that are combined into a complex, iterative workflow in which workers check and improve each other's results. This paper raises an exciting question for AI — could an autonomous agent control these workflows without human intervention, yielding better results than today's state of the art, a fixed control program? We describe a planner, TurKontrol, that formulates workflow control as a decision-theoretic optimization problem, trading off the implicit quality of a solution artifact against the cost for workers to achieve it. We lay the mathematical framework to govern the various decisions at each point in a popular class of workflows. Based on our analysis we implement the workflow control algorithm and present experiments demonstrating that TurKontrol obtains much higher utilities than popular fixed policies. Peng Dai 0001, Mausam, Daniel S. Weld |
AAAI | 2 |
| 2010 | SixthSense: Fast and Reliable Recognition of Dead Ends in MDPsabstractThe results of the latest International Probabilistic Planning Competition (IPPC-2008) indicate that the presence of dead ends, states with no trajectory to the goal, makes MDPs hard for modern probabilistic planners. Implicit dead ends, states with executable actions but no path to the goal, are particularly challenging; existing MDP solvers spend much time and memory identifying these states. As a first attempt to address this issue, we propose a machine learning algorithm called SIXTHSENSE. SIXTHSENSE helps existing MDP solvers by finding nogoods, conjunctions of literals whose truth in a state implies that the state is a dead end. Importantly, our learned nogoods are sound, and hence the states they identify are true dead ends. SIXTHSENSE is very fast, needs little training data, and takes only a small fraction of total planning time. While IPPC problems may have millions of dead ends, they may typically be represented with only a dozen or two no-goods. Thus, nogood learning efficiently produces a quick and reliable means for dead-end recognition. Our experiments show that the nogoods found by SIXTHSENSE routinely reduce planning space and time on IPPC domains, enabling some planners to solve problems they could not previously handle. Andrey Kolobov, Mausam, Daniel S. Weld |
AAAI | 2 |
| 2010 | Panlingual Lexical Translation via Probabilistic InferenceabstractThe bare minimum lexical resource required to translate between a pair of languages is a translation dictionary. Unfortunately, dictionaries exist only between a tiny fraction of the 49 million possible language-pairs making machine translation virtually impossible between most of the languages. This paper summarizes the last four years of our research motivated by the vision of panlingual communication. Our research comprises three key steps. First, we compile over 630 freely available dictionaries over the Web and convert this data into a single representation – the translation graph. Second, we build several inference algorithms that infer translations between word pairs even when no dictionary lists them as translations. Finally, we run our inference procedure offline to construct PANDICTIONARY– a sense-distinguished, massively multilingual dictionary that has translations in more than 1000 languages. Our experiments assess the quality of this dictionary and find that we have 4 times as many translations at a high precision of 0.9 compared to the English Wiktionary, which is the lexical resource closest to PANDICTIONARY. Mausam, Stephen Soderland, Oren Etzioni |
AAAI | 1 |
| 2010 | A Latent Dirichlet Allocation Method for Selectional Preferences
Alan Ritter, Mausam, Oren Etzioni |
ACL | 2 |
| 2010 | Identifying Functional Relations in Web Text
Thomas Lin, Mausam, Oren Etzioni |
EMNLP | 2 |
| 2010 | Panlingual lexical translation via probabilistic inference
Mausam, Stephen Soderland, Oren Etzioni, Daniel S. Weld, Kobi Reiter, Michael Skinner, Marcus Sammer, Jeff A. Bilmes |
Artif. Intell. | 1 |
| 2009 | Compiling a Massive, Multilingual Dictionary via Probabilistic Inference
Mausam, Stephen Soderland, Oren Etzioni, Daniel S. Weld, Michael Skinner, Jeff A. Bilmes |
ACL/IJCNLP | 1 |
| 2009 | Domain-Independent, Automatic Partitioning for Probabilistic Planning
Peng Dai 0001, Mausam, Daniel S. Weld |
IJCAI | 2 |
| 2009 | ReTrASE: Integrating Paradigms for Approximate Probabilistic Planning
Andrey Kolobov, Mausam, Daniel S. Weld |
IJCAI | 2 |
| 2009 | Lemmatic Machine Translation
Stephen Soderland, Christopher Lim, Mausam, Oren Etzioni, Jonathan Pool |
MTSummit | 3 |
| 2009 | A Heuristic Search Approach to Planning with Continuous Resources in Stochastic DomainsabstractWe consider the problem of optimal planning in stochastic domains with resource constraints, where the resources are continuous and the choice of action at each step depends on resource availability. We introduce the HAO* algorithm, a generalization of the AO* algorithm that performs search in a hybrid state space that is modeled using both discrete and continuous state variables, where the continuous variables represent monotonic resources. Like other heuristic search algorithms, HAO* leverages knowledge of the start state and an admissible heuristic to focus computational effort on those parts of the state space that could be reached from the start state by following an optimal policy. We show that this approach is especially effective when resource constraints limit how much of the state space is reachable. Experimental results demonstrate its effectiveness in the domain that motivates our research: automated planning for planetary exploration rovers. Nicolas Meuleau, Emmanuel Benazera, Ronen I. Brafman, Eric A. Hansen, Mausam |
J. Artif. Intell. Res. | 5 |
| 2008 | Partitioned External-Memory Value Iteration
Peng Dai 0001, Mausam, Daniel S. Weld |
AAAI | 2 |
| 2008 | Planning with Durative Actions in Stochastic DomainsabstractProbabilistic planning problems are typically modeled as a Markov Decision Process (MDP). MDPs, while an otherwise expressive model, allow only for sequential, non-durative actions. This poses severe restrictions in modeling and solving a real world planning problem. We extend the MDP model to incorporate 1) simultaneous action execution, 2) durative actions, and 3) stochastic durations. We develop several algorithms to combat the computational explosion introduced by these features. The key theoretical ideas used in building these algorithms are -- modeling a complex problem as an MDP in extended state/action space, pruning of irrelevant actions, sampling of relevant actions, using informed heuristics to guide the search, hybridizing different planners to achieve benefits of both, approximating the problem and replanning. Our empirical evaluation illuminates the different merits in using various algorithms, viz., optimality, empirical closeness to optimality, theoretical error bounds, and speed. Mausam, Daniel S. Weld |
J. Artif. Intell. Res. | 1 |
| 2007 | When is Temporal Planning Really Temporal?
William Cushing, Subbarao Kambhampati, Mausam, Daniel S. Weld |
IJCAI | 3 |
| 2007 | A Hybridized Planner for Stochastic Domains
Mausam, Piergiorgio Bertoli, Daniel S. Weld |
IJCAI | 1 |
| 2006 | Probabilistic Temporal Planning with Uncertain Durations
Mausam, Daniel S. Weld |
AAAI | 1 |
| 2005 | Planning with Continuous Resources in Stochastic Domains
Mausam, Emmanuel Benazera, Ronen I. Brafman, Nicolas Meuleau, Eric A. Hansen |
IJCAI | 1 |
| 2004 | Solving Concurrent Markov Decision Processes
Mausam, Daniel S. Weld |
AAAI | 1 |
| 2004 | Adversarial classificationabstractEssentially all data mining algorithms assume that the data-generating process is independent of the data miner's activities. However, in many domains, including spam detection, intrusion detection, fraud detection, surveillance and counter-terrorism, this is far from the case: the data is actively manipulated by an adversary seeking to make the classifier produce false negatives. In these domains, the performance of a classifier can degrade rapidly after it is deployed, as the adversary learns to defeat it. Currently the only solution to this is repeated, manual, ad hoc reconstruction of the classifier. In this paper we develop a formal framework and algorithms for this problem. We view classification as a game between the classifier and the adversary, and produce a classifier that is optimal given the adversary's optimal strategy. Experiments in a spam detection domain show that this approach can greatly outperform a classifier learned in the standard way, and (within the parameters of the problem) automatically adapt the classifier to the adversary's evolving manipulations. Nilesh N. Dalvi, Pedro M. Domingos, Mausam, Sumit K. Sanghai |
KDD | 3 |