EDBT 2026 Demo / reviewers in the wild / expert
Niket Tandon
dblp:29/9923
· DBLP profile ↗
36ranked-venue papers
9as first author
16since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 9 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Calibrating Large Language Models with Sample ConsistencyabstractAccurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we derive model confidence from the distribution of multiple randomly sampled generations, using three measures of consistency. We extensively evaluate eleven open and closed-source models on nine reasoning datasets. Results show that consistency-based calibration methods outperform existing post-hoc approaches in terms of calibration error. Meanwhile, we find that factors such as intermediate explanations, model scaling, and larger sample sizes enhance calibration, while instruction-tuning makes calibration more difficult. Moreover, confidence scores obtained from consistency can potentially enhance model performance. Finally, we offer guidance on choosing suitable consistency metrics for calibration, tailored to model characteristics such as the exposure to instruction-tuning and RLHF. Qing Lyu 0001, Kumar Shridhar, Chaitanya Malaviya, Li Zhang 0039, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, Chris Callison-Burch |
AAAI | 6 |
| 2025 | On the Reliability of Large Language Models for Causal DiscoveryabstractThis study investigates the efficacy of Large Language Models (LLMs) in causal discovery. Using newly available open-source LLMs, OLMo and BLOOM, which provide access to their pre-training corpora, we investigate how LLMs address causal discovery through three research questions. We examine: (i) the impact of memorization for accurate causal relation prediction, (ii) the influence of incorrect causal relations in pre-training data, and (iii) the contextual nuances that influence LLMs’ understanding of causal relations. Our findings indicate that while LLMs are effective in recognizing causal relations that occur frequently in pre-training data, their ability to generalize to new or rare causal relations is limited. Moreover, the presence of incorrect causal relations significantly undermines the confidence of LLMs in corresponding correct causal relations, and the contextual information critically affects the outcomes of LLMs to discern causal connections between random variables. Tao Feng 0013, Lizhen Qu, Niket Tandon, Zhuang Li 0001, Xiaoxi Kang, Gholamreza Haffari |
ACL (1) | 3 |
| 2025 | IRIS: An Iterative and Integrated Framework for Verifiable Causal Discovery in the Absence of Tabular DataabstractCausal discovery is fundamental to scientific research, yet traditional statistical algorithms face significant challenges, including expensive data collection, redundant computation for known relations, and unrealistic assumptions.While recent LLM-based methods excel at identifying commonly known causal relations, they fail to uncover novel relations.We introduce IRIS (Iterative Retrieval and Integrated System for Real-Time Causal Discovery), a novel framework that addresses these limitations.Starting with a set of initial variables, IRIS automatically collects relevant documents, extracts variables, and uncovers causal relations.Our hybrid causal discovery method combines statistical algorithms and LLM-based methods to discover known and novel causal relations.In addition to causal discovery on initial variables, the missing variable proposal component of IRIS identifies and incorporates missing variables to expand the causal graphs.Our approach enables real-time causal discovery from only a set of initial variables without requiring pre-existing datasets. Tao Feng 0013, Lizhen Qu, Niket Tandon, Gholamreza Haffari |
ACL (1) | 3 |
| 2025 | MOGIC: Metadata-infused Oracle Guidance for Improved Extreme ClassificationabstractRetrieval-augmented classification and generation models benefit from *early-stage fusion* of high-quality text-based metadata, often called memory, but face high latency and noise sensitivity. In extreme classification (XC), where low latency is crucial, existing methods use *late-stage fusion* for efficiency and robustness. To enhance accuracy while maintaining low latency, we propose MOGIC, a novel approach to metadata-infused oracle guidance for XC. We train an early-fusion oracle classifier with access to both query-side and label-side ground-truth metadata in textual form and subsequently use it to guide existing memory-based XC disciple models via regularization. The MOGIC algorithm improves precision@1 and propensity-scored precision@1 of XC disciple models by 1-2% on six standard datasets, at no additional inference-time cost. We show that MOGIC can be used in a plug-and-play manner to enhance memory-free XC models such as NGAME or DEXA. Lastly, we demonstrate the robustness of the MOGIC algorithm to missing and noisy metadata. The code is publicly available at [https://github.com/suchith720/mogic](https://github.com/suchith720/mogic). Suchith C. Prabhu, Bhavyajeet Singh, Anshul Mittal, Siddarth Asokan, Shikhar Mohan, Deepak Saini, Yashoteja Prabhu, Lakshya Kumar, Jian Jiao 0007, Amit Singh 0003, Niket Tandon, Sumeet Agarwal, Manik Varma |
ICML | 11 |
| 2024 | WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language ModelsabstractThe awareness of multi-cultural human values is critical to the ability of language models (LMs) to generate safe and personalized responses. However, this awareness of LMs has been insufficiently studied, since the computer science community lacks access to the large-scale real-world data about multi-cultural values. In this paper, we present WorldValuesBench, a globally diverse, large-scale benchmark dataset for the multi-cultural value prediction task, which requires a model to generate a rating response to a value question based on demographic contexts. Our dataset is derived from an influential social science project, World Values Survey (WVS), that has collected answers to hundreds of value questions (e.g., social, economic, ethical) from 94,728 participants worldwide. We have constructed more than 20 million examples of the type "(demographic attributes, value question) → answer” from the WVS responses. We perform a case study using our dataset and show that the task is challenging for strong open and closed-source models. On merely 11.1%, 25.0%, 72.2%, and 75.0% of the questions, Alpaca-7B, Vicuna-7B-v1.5, Mixtral-8x7B-Instruct-v0.1, and GPT-3.5 Turbo can respectively achieve <0.2 Wasserstein 1-distance from the human normalized answer distributions. WorldValuesBench opens up new research avenues in studying limitations and opportunities in multi-cultural value awareness of LMs. Wenlong Zhao 0001, Debanjan Mondal, Niket Tandon, Danica Dillion, Kurt Gray, Yuling Gu |
LREC/COLING | 3 |
| 2024 | OpenPI2.0: An Improved Dataset for Entity Tracking in TextsabstractLi Zhang, Hainiu Xu, Abhinav Kommula, Chris Callison-Burch, Niket Tandon. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Li Zhang 0039, Hainiu Xu, Abhinav Kommula, Chris Callison-Burch, Niket Tandon |
EACL (1) | 5 |
| 2024 | Let Me Teach You: Pedagogical Foundations of Feedback for Language ModelsabstractNatural Language Feedback (NLF) is an increasingly popular mechanism for aligning Large Language Models (LLMs) to human preferences.Despite the diversity of the information it can convey, NLF methods are often handdesigned and arbitrary, with little systematic grounding.At the same time, research in learning sciences has long established several effective feedback models.In this opinion piece, we compile ideas from pedagogy to introduce FELT, a feedback framework for LLMs that outlines various characteristics of the feedback space, and a feedback content taxonomy based on these variables, providing a general mapping of the feedback space.In addition to streamlining NLF designs, FELT also brings out new, unexplored directions for research in NLF.We make our taxonomy available to the community, providing guides and examples for mapping our categorizations to future research. Beatriz Borges, Niket Tandon, Tanja Käser, Antoine Bosselut |
EMNLP | 2 |
| 2024 | In-Context Principle Learning from MistakesabstractIn-context learning (ICL, also known as few-shot prompting) has been the standard method of adapting LLMs to downstream tasks, by learning from a few input-output examples. Nonetheless, all ICL-based approaches only learn from correct input-output pairs. In this paper, we revisit this paradigm, by learning more from the few given input-output examples. We introduce Learning Principles (LEAP): First, we intentionally induce the model to make mistakes on these few examples; then we reflect on these mistakes, and learn explicit task-specific “principles” from them, which help solve similar problems and avoid common mistakes; finally, we prompt the model to answer unseen test questions using the original few-shot examples and these learned general principles. We evaluate LEAP on a wide range of benchmarks, including multi-hop question answering (Hotpot QA), textual QA (DROP), Big-Bench Hard reasoning, and math problems (GSM8K and MATH); in all these benchmarks, LEAP improves the strongest available LLMs such as GPT-3.5-turbo, GPT-4, GPT-4-turbo and Claude-2.1. For example, LEAP improves over the standard few-shot prompting using GPT-4 by 7.5% in DROP, and by 3.3% in HotpotQA. Importantly, LEAP does not require any more input or examples than the standard few-shot prompting settings. Tianjun Zhang, Aman Madaan, Luyu Gao, Steven Zheng, Swaroop Mishra, Yiming Yang 0002, Niket Tandon, Uri Alon 0002 |
ICML | 7 |
| 2023 | RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model OutputsabstractAfra Feyza Akyurek, Ekin Akyurek, Ashwin Kalyan, Peter Clark, Derry Tanti Wijaya, Niket Tandon. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Afra Feyza Akyürek, Ekin Akyürek, Ashwin Kalyan, Peter Clark, Derry Wijaya, Niket Tandon |
ACL (1) | 6 |
| 2023 | Editing Common Sense in TransformersabstractEditing model parameters directly in Transformers makes updating open-source transformer-based models possible without re-training (Meng et al., 2023).However, these editing methods have only been evaluated on statements about encyclopedic knowledge with a single correct answer.Commonsense knowledge with multiple correct answers, e.g., an apple can be green or red but not transparent, has not been studied but is as essential for enhancing transformers' reliability and usefulness.In this paper, we investigate whether commonsense judgments are causally associated with localized, editable parameters in Transformers, and we provide an affirmative answer.We find that directly applying the MEMIT editing algorithm results in sub-par performance, and propose to improve it for the commonsense domain by varying edit tokens and improving the layer selection strategy, i.e., MEMIT CSK .GPT-2 Large and XL models edited using MEMIT CSK outperform best-fine-tuned baselines by 10.97% and 10.73% F1 scores on PEP3k and 20Q datasets.In addition, we propose a novel evaluation dataset, PROBE SET, that contains unaffected and affected neighborhoods, affected paraphrases, and affected reasoning challenges.MEMIT CSK performs well across the metrics while fine-tuning baselines show significant trade-offs between unaffected and affected metrics.These results suggest a compelling future direction for incorporating feedback about common sense into Transformers through direct model editing. 1 * Co-first and last authors.Lorraine's work done at AI2. 1 Code and datasets for all experiments are available at https://github.com/anshitag/memit_csk Anshita Gupta, Debanjan Mondal, Akshay Krishna Sheshadri, Wenlong Zhao 0001, Xiang Li 0069, Sarah Wiegreffe, Niket Tandon |
EMNLP | 7 |
| 2023 | Self-Refine: Iterative Refinement with Self-FeedbackabstractLike humans, large language models (LLMs) do not always generate the best output on their first try. Motivated by how humans refine their written text, we introduce Self-Refine, an approach for improving initial outputs from LLMs through iterative feedback and refinement. The main idea is to generate an initial output using an LLMs; then, the same LLMs provides *feedback* for its output and uses it to *refine* itself, iteratively. Self-Refine does not require any supervised training data, additional training, or reinforcement learning, and instead uses a single LLM as the generator, refiner and the feedback provider. We evaluate Self-Refine across 7 diverse tasks, ranging from dialog response generation to mathematical reasoning, using state-of-the-art (GPT-3.5, ChatGPT, and GPT-4) LLMs. Across all evaluated tasks, outputs generated with Self-Refine are preferred by humans and automatic metrics over those generated with the same LLM using conventional one-step generation, improving by $\sim$20\% absolute on average in task performance. Our work demonstrates that even state-of-the-art LLMs like GPT-4 can be further improved at test-time using our simple, standalone approach. Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon 0002, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang 0002, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, Peter Clark |
NeurIPS | 2 |
| 2022 | Using Commonsense Knowledge to Answer Why-QuestionsabstractYash Kumar Lal, Niket Tandon, Tanvi Aggarwal, Horace Liu, Nathanael Chambers, Raymond Mooney, Niranjan Balasubramanian. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Yash Kumar Lal, Niket Tandon, Tanvi Aggarwal, Horace Liu, Nathanael Chambers, Raymond J. Mooney, Niranjan Balasubramanian |
EMNLP | 2 |
| 2022 | Conditional set generation using Seq2seq modelsabstractConditional set generation learns a mapping from an input sequence of tokens to a set.Several NLP tasks, such as entity typing and dialogue emotion tagging, are instances of set generation.SEQ2SEQ models, a popular choice for set generation, treat a set as a sequence and do not fully leverage its key properties, namely order-invariance and cardinality.We propose a novel algorithm for effectively sampling informative orders over the combinatorial space of label orders.We jointly model the set cardinality and output by prepending the set size and taking advantage of the autoregressive factorization used by SEQ2SEQ models.Our method is a model-independent data augmentation approach that endows any SEQ2SEQ model with the signals of order-invariance and cardinality.Training a SEQ2SEQ model on this augmented data (without any additional annotations) gets an average relative improvement of 20% on four benchmark datasets across various models: BART-base, T5-11B, and GPT3-175B. 1 Aman Madaan, Dheeraj Rajagopal, Niket Tandon, Yiming Yang 0002, Antoine Bosselut |
EMNLP | 3 |
| 2022 | Memory-assisted prompt editing to improve GPT-3 after deploymentabstractLarge LMs such as GPT-3 are powerful, but can commit mistakes that are obvious to humans.For example, GPT-3 would mistakenly interpret "What word is similar to good?" to mean a homophone, while the user intended a synonym.Our goal is to effectively correct such errors via user interactions with the system but without retraining, which will be prohibitively costly.We pair GPT-3 with a growing memory of recorded cases where the model misunderstood the user's intents, along with user feedback for clarification.Such a memory allows our system to produce enhanced prompts for any new query based on the user feedback for error correction on similar cases in the past.On four tasks (two lexical tasks, two advanced ethical reasoning tasks), we show how a (simulated) user can interactively teach a deployed GPT-3, substantially increasing its accuracy over the queries with different kinds of misunderstandings by the GPT-3.Our approach is a step towards the low-cost utility enhancement for very large pre-trained LMs. 1 * Equal Contribution 1 Code, data, and instructions to implement MemPrompt for a new task at https://www.memprompt.com/ Aman Madaan, Niket Tandon, Peter Clark, Yiming Yang 0002 |
EMNLP | 2 |
| 2021 | Think about it! Improving defeasible reasoning by first modeling the question scenarioabstractDefeasible reasoning is the mode of reasoning where conclusions can be overturned by taking into account new evidence.Existing cognitive science literature on defeasible reasoning suggests that a person forms a mental model of the problem scenario before answering questions.Our research goal asks whether neural models can similarly benefit from envisioning the question scenario before answering a defeasible query.Our approach is, given a question, to have a model first create a graph of relevant influences, and then leverage that graph as an additional input when answering the question.Our system, CURIOUS, achieves a new stateof-the-art on three different defeasible reasoning datasets.This result is significant as it illustrates that performance can be improved by guiding a system to "think about" a question and explicitly model the scenario, rather than answering reflexively. 1 Aman Madaan, Niket Tandon, Dheeraj Rajagopal, Peter Clark, Yiming Yang 0002, Eduard H. Hovy |
EMNLP (1) | 2 |
| 2021 | Information to Wisdom: Commonsense Knowledge Extraction and CompilationabstractCommonsense knowledge is a foundational cornerstone of artificial intelligence applications. Whereas information extraction and knowledge base construction for instance-oriented assertions, such as Brad Pitt's birth date, or Angelina Jolie's movie awards, has received much attention, commonsense knowledge on general concepts (politicians, bicycles, printers) and activities (eating pizza, fixing printers) has only been tackled recently. In this tutorial we present state-of-the-art methodologies towards the compilation and consolidation of such commonsense knowledge (CSK). We cover text-extraction-based, multi-modal and Transformer-based techniques, with special focus on the issues of web search and ranking, as of relevance to the WSDM community. Simon Razniewski, Niket Tandon, Aparna S. Varde |
WSDM | 2 |
| 2020 | I Am Guessing You Can't Recognize This: Generating Adversarial Images for Object Detection Using Spatial Commonsense (Student Abstract)abstractCan we automatically predict failures of an object detection model on images from a target domain? We characterize errors of a state-of-the-art object detection model on the currently popular smart mobility domain, and find that a large number of errors can be identified using spatial commonsense. We propose øurmodel , a system that automatically identifies a large number of such errors based on commonsense knowledge. Our system does not require any new annotations and can still find object detection errors with high accuracy (more than 80% when measured by humans). This work lays the foundation to answer exciting research questions on domain adaptation including the ability to automatically create adversarial datasets for target domain. Anurag Garg, Niket Tandon, Aparna S. Varde |
AAAI | 2 |
| 2020 | A Dataset for Tracking Entities in Open Domain Procedural TextabstractNiket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson, Eduard Hovy. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Niket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson 0001, Eduard H. Hovy |
EMNLP (1) | 1 |
| 2020 | Using Commonsense Knowledge and Text Mining for Implicit Requirements LocalizationabstractThis paper addresses identification of implicit requirements (IMRs) in software requirements specifications (SRS). IMRs, as opposed to explicit requirements, are not specified by users but are more subtle. It has been noticed that IMRs are crucial to the success of software development. In this paper, we demonstrate a software tool called COTIR developed by us as a system that integrates Commonsense knowledge, Ontology and Text mining for early identification of Implicit Requirements. This relieves human software engineers from the tedious task of manually identifying IMRs in huge SRS documents. Our evaluation reveals that COTIR outperforms existing IMR tools. This demo paper would be useful to Software Engineers since it deals with automation in the requirements analysis phase, thus contributing to Requirements Engineering. It would interest AI scientists as it entails multi-disciplinary work encompassing text mining, ontology and commonsense knowledge. It makes a broader impact on Smart Cities, because automated identification of IMRs would offer inputs to Smart City Tools, where requirements may often be implicit given that Smart Cities are an emerging and growing paradigm. Onyeka Emebo, Aparna S. Varde, Vaibhav K. Anu, Niket Tandon, Olawande J. Daramola |
ICTAI | 4 |
| 2019 | Everything Happens for a Reason: Discovering the Purpose of Actions in Procedural TextabstractBhavana Dalvi, Niket Tandon, Antoine Bosselut, Wen-tau Yih, Peter Clark. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Bhavana Dalvi, Niket Tandon, Antoine Bosselut, Scott Yih, Peter Clark |
EMNLP/IJCNLP (1) | 2 |
| 2019 | WIQA: A dataset for "What if..." reasoning over procedural textabstractNiket Tandon, Bhavana Dalvi, Keisuke Sakaguchi, Peter Clark, Antoine Bosselut. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Niket Tandon, Bhavana Dalvi, Keisuke Sakaguchi, Peter Clark, Antoine Bosselut |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Reasoning about Actions and State Changes by Injecting Commonsense KnowledgeabstractComprehending procedural text, e.g., a paragraph describing photosynthesis, requires modeling actions and the state changes they produce, so that questions about entities at different timepoints can be answered.Although several recent systems have shown impressive progress in this task, their predictions can be globally inconsistent or highly improbable.In this paper, we show how the predicted effects of actions in the context of a paragraph can be improved in two ways: (1) by incorporating global, commonsense constraints (e.g., a non-existent entity cannot be destroyed), and (2) by biasing reading with preferences from large-scale corpora (e.g., trees rarely move).Unlike earlier methods, we treat the problem as a neural structured prediction task, allowing hard and soft constraints to steer the model away from unlikely predictions.We show that the new model significantly outperforms earlier systems on a benchmark dataset for procedural text comprehension (+8% relative gain), and that it also avoids some of the nonsensical predictions that earlier systems make. Niket Tandon, Bhavana Dalvi, Joel Grus, Scott Yih, Antoine Bosselut, Peter Clark |
EMNLP | 1 |
| 2018 | Tracking State Changes in Procedural Text: a Challenge Dataset and Models for Process Paragraph ComprehensionabstractBhavana Dalvi, Lifu Huang, Niket Tandon, Wen-tau Yih, Peter Clark. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Bhavana Dalvi, Lifu Huang, Niket Tandon, Scott Yih, Peter Clark |
NAACL-HLT | 3 |
| 2018 | VISIR: Visual and Semantic Image Label RefinementabstractThe social media explosion has populated the Internet with a wealth of images. There are two existing paradigms for image retrieval: 1)content-based image retrieval (BIR), which has traditionally used visual features for similarity search (e.g., SIFT features), and 2) tag-based image retrieval (TBIR), which has relied on user tagging (e.g., Flickr tags). CBIR now gains semantic expressiveness by advances in deep-learning-based detection of visual labels. TBIR benefits from query-and-click logs to automatically infer more informative labels. However, learning-based tagging still yields noisy labels and is restricted to concrete objects, missing out on generalizations and abstractions. Click-based tagging is limited to terms that appear in the textual context of an image or in queries that lead to a click. This paper addresses the above limitations by semantically refining and expanding the labels suggested by learning-based object detection. We consider the semantic coherence between the labels for different objects, leverage lexical and commonsense knowledge, and cast the label assignment into a constrained optimization problem solved by an integer linear program. Experiments show that our method, called VISIR, improves the quality of the state-of-the-art visual labeling tools like LSDA and YOLO. Sreyasi Nag Chowdhury, Niket Tandon, Hakan Ferhatosmanoglu, Gerhard Weikum |
WSDM | 2 |
| 2017 | Distilling Task Knowledge from How-To CommunitiesabstractKnowledge graphs have become a fundamental asset for search engines. A fair amount of user queries seek information on problem-solving tasks such as building a fence or repairing a bicycle. However, knowledge graphs completely lack this kind of how-to knowledge. This paper presents a method for automatically constructing a formal knowledge base on tasks and task-solving steps, by tapping the contents of online communities such as WikiHow. We employ Open-IE techniques to extract noisy candidates for tasks, steps and the required tools and other items. For cleaning and properly organizing this data, we devise embedding-based clustering techniques. The resulting knowledge base, HowToKB, includes a hierarchical taxonomy of disambiguated tasks, temporal orders of sub-tasks, and attributes for involved items. A comprehensive evaluation of HowToKB shows high accuracy. As an extrinsic use case, we evaluate automatically searching related YouTube videos for HowToKB tasks. Cuong Xuan Chu, Niket Tandon, Gerhard Weikum |
WWW | 2 |
| 2017 | Movie DescriptionabstractAudio description (AD) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by design mainly visual and thus naturally form an interesting data source for computer vision and computational linguistics. In this work we propose a novel dataset which contains transcribed ADs, which are temporally aligned to full length movies. In addition we also collected and aligned movie scripts used in prior work and compare the two sources of descriptions. We introduce the Large Scale Movie Description Challenge (LSMDC) which contains a parallel corpus of 128,118 sentences aligned to video clips from 200 movies (around 150 h of video in total). The goal of the challenge is to automatically generate descriptions for the movie clips. First we characterize the dataset by benchmarking different approaches for generating video descriptions. Comparing ADs to scripts, we find that ADs are more visual and describe precisely what is shown rather than what should happen according to the scripts created prior to movie production. Furthermore, we present and compare the results of several teams who participated in the challenges organized in the context of two workshops at ICCV 2015 and ECCV 2016. Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Joseph Pal, Hugo Larochelle, Aaron C. Courville, Bernt Schiele |
Int. J. Comput. Vis. | 4 |
| 2017 | Domain-Targeted, High Precision Knowledge ExtractionabstractOur goal is to construct a domain-targeted, high precision knowledge base (KB), containing general (subject,predicate,object) statements about the world, in support of a downstream question-answering (QA) application. Despite recent advances in information extraction (IE) techniques, no suitable resource for our task already exists; existing resources are either too noisy, too named-entity centric, or too incomplete, and typically have not been constructed with a clear scope or purpose. To address these, we have created a domain-targeted, high precision knowledge extraction pipeline, leveraging Open IE, crowdsourcing, and a novel canonical schema learning algorithm (called CASI), that produces high precision knowledge targeted to a particular domain - in our case, elementary science. To measure the KB’s coverage of the target domain’s knowledge (its “comprehensiveness” with respect to science) we measure recall with respect to an independent corpus of domain text, and show that our pipeline produces output with over 80% precision and 23% recall with respect to that target, a substantially higher coverage of tuple-expressible science knowledge than other comparable resources. We have made the KB publicly available. Bhavana Dalvi, Niket Tandon, Peter Clark |
Trans. Assoc. Comput. Linguistics | 2 |
| 2016 | Commonsense in Parts: Mining Part-Whole Relations from the Web and Image TagsabstractCommonsense knowledge about part-whole relations (e.g., screen partOf notebook) is important for interpreting user input in web search and question answering, or for object detection in images. Prior work on knowledge base construction has compiled part-whole assertions, but with substantial limitations: i) semantically different kinds of part-whole relations are conflated into a single generic relation, ii) the arguments of a part-whole assertion are merely words with ambiguous meaning, iii) the assertions lack additional attributes like visibility (e.g., a nose is visible but a kidney is not) and cardinality information (e.g., a bird has two legs while a spider eight), iv) limited coverage of only tens of thousands of assertions. This paper presents a new method for automatically acquiring part-whole commonsense from Web contents and image tags at an unprecedented scale, yielding many millions of assertions, while specifically addressing the four shortcomings of prior work. Our method combines pattern-based information extraction methods with logical reasoning. We carefully distinguish different relations: physicalPartOf, memberOf, substanceOf. We consistently map the arguments of all assertions onto WordNet senses, eliminating the ambiguity of word-level assertions. We identify whether the parts can be visually perceived, and infer cardinalities for the assertions. The resulting commonsense knowledge base has very high quality and high coverage, with an accuracy of 89% determined by extensive sampling, and is publicly available. Niket Tandon, Charles Hariman, Jacopo Urbani, Anna Rohrbach, Marcus Rohrbach, Gerhard Weikum |
AAAI | 1 |
| 2016 | WebBrain: Joint Neural Learning of Large-Scale Commonsense Knowledge
Niket Tandon, Charles Hariman, Gerard de Melo |
ISWC (1) | 2 |
| 2015 | Multimedia Data for the Visually ImpairedabstractThe Web contains a large amount of information in the form of videos that remains inaccessible to the visually impaired people. We identify a class of videos whose information content can be approximately encoded as an audio, thereby increasing the amount of accessible videos. We propose a model to automatically identify such videos. Our model jointly relies on the textual metadata and visual content of the video. We use this model to re-rank Youtube video search results based on accessibility of the video. We present preliminary results by conducting a user study with visually impaired people to measure the effectiveness of our system. Niket Tandon, Shekhar Sharma, Tanima Makkad |
AAAI | 1 |
| 2015 | Perceptually Grounded Selectional PreferencesabstractEkaterina Shutova, Niket Tandon, Gerard de Melo. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Ekaterina Shutova, Niket Tandon, Gerard de Melo |
ACL (1) | 2 |
| 2015 | Knowlywood: Mining Activity Knowledge From Hollywood NarrativesabstractDespite the success of large knowledge bases, one kind of knowledge that has not received attention so far is that of human activities. An example of such an activity is proposing to someone (to get married). For the computer, knowing that this involves two adults, often but not necessarily a woman and a man, that it often takes place in some romantic location, that it typically involves flowers or jewelry, and that it is usually followed by kissing, is a valuable asset for tasks like natural language dialog, scene understanding, or video search. Niket Tandon, Gerard de Melo, Abir De, Gerhard Weikum |
CIKM | 1 |
| 2015 | A dataset for Movie DescriptionabstractAudio Description (AD) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by design mainly visual and thus naturally form an interesting data source for computer vision and computational linguistics. In this work we propose a novel dataset which contains transcribed ADs, which are temporally aligned to full length HD movies. In addition we also collected the aligned movie scripts which have been used in prior work and compare the two different sources of descriptions. In total the MPII Movie Description dataset (MPII-MD) contains a parallel corpus of over 68K sentences and video snippets from 94 HD movies. We characterize the dataset by benchmarking different approaches for generating video descriptions. Comparing ADs to scripts, we find that ADs are far more visual and describe precisely what is shown rather than what should happen according to the scripts created prior to movie production. Anna Rohrbach, Marcus Rohrbach, Niket Tandon, Bernt Schiele |
CVPR | 3 |
| 2014 | Acquiring Comparative Commonsense Knowledge from the WebabstractApplications are increasingly expected to make smart decisions based on what humans consider basic commonsense. An often overlooked but essential form of commonsense involves comparisons, e.g. the fact that bears are typically more dangerous than dogs, that tables are heavier than chairs, or that ice is colder than water. In this paper, we first rely on open information extraction methods to obtain large amounts of comparisons from the Web. We then develop a joint optimization model for cleaning and disambiguating this knowledge with respect to WordNet. This model relies on integer linear programming and semantic coherence scores. Experiments show that our model outperforms strong baselines and allows us to obtain a large knowledge base of disambiguated commonsense assertions. Niket Tandon, Gerard de Melo, Gerhard Weikum |
AAAI | 1 |
| 2014 | WebChild: harvesting and organizing commonsense knowledge from the webabstractThis paper presents a method for automatically constructing a large commonsense knowledge base, called WebChild, from Web contents. WebChild contains triples that connect nouns with adjectives via fine-grained relations like hasShape, hasTaste, evokesEmotion, etc. The arguments of these assertions, nouns and adjectives, are disambiguated by mapping them onto their proper WordNet senses. Our method is based on semi-supervised Label Propagation over graphs of noisy candidate assertions. We automatically derive seeds from WordNet and by pattern matching from Web text collections. The Label Propagation algorithm provides us with domain sets and range sets for 19 different relations, and with confidence-ranked assertions between WordNet senses. Large-scale experiments demonstrate the high accuracy (more than 80 percent) and coverage (more than four million fine grained disambiguated assertions) of WebChild. Niket Tandon, Gerard de Melo, Fabian M. Suchanek, Gerhard Weikum |
WSDM | 1 |
| 2011 | Deriving a Web-Scale Common Sense Fact DatabaseabstractThe fact that birds have feathers and ice is cold seems trivially true. Yet, most machine-readable sources of knowledge either lack such common sense facts entirely or have only limited coverage. Prior work on automated knowledge base construction has largely focused on relations between named entities and on taxonomic knowledge, while disregarding common sense properties. In this paper, we show how to gather large amounts of common sense facts from Web n-gram data, using seeds from the ConceptNet collection. Our novel contributions include scalable methods for tapping onto Web-scale data and a new scoring model to determine which patterns and facts are most reliable. The experimental results show that this approach extends ConceptNet by many orders of magnitude at comparable levels of precision. Niket Tandon, Gerard de Melo, Gerhard Weikum |
AAAI | 1 |