Chitta Baral

dblp:b/ChittaBaral · DBLP profile ↗
← Back
160ranked-venue papers
60as first author
49since 2021 · last 2026
0000-0002-7549-723XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 126 · 47 first-author · 44 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 18 first-author · 11 since 2021Theory of computation · 29 · 21 first-authorSoftware engineering, systems software and programming languages · 10 · 7 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-authorSystems, architecture and hardware · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
abstract
Recent advances in multimodal large language models (MLLMs) have yielded increasingly powerful models, yet their perceptual capacities remain poorly characterized. In practice, most model families scale language component while reusing nearly identical vision encoders (e.g., Qwen2.5-VL 3B/7B/72B), which raises pivotal concerns about whether progress reflects genuine visual grounding or reliance on internet-scale textual world knowledge. Existing evaluation methods emphasize end-task accuracy, overlooking robustness, attribution fidelity, and reasoning under controlled perturbations. We present The Perceptual Observatory, a framework that characterizes MLLMs across verticals like: (i) simple vision tasks, such as face matching and text-in-vision comprehension capabilities; (ii) local-to-global understanding, encompassing image matching, grid pointing game, and attribute localization, which tests general visual grounding. Each vertical is instantiated with ground-truth datasets of faces and words, systematically perturbed through pixel-based augmentations and diffusion-based stylized illusions. The Perceptual Observatory moves beyond leaderboard accuracy to yield insights into how MLLMs preserve perceptual grounding and relational structure under perturbations, providing a principled foundation for analyzing strengths and weaknesses of current and future models.
Tejas Anvekar, Fenil Denish Bardoliya, Pavan Turaga, Chitta Baral, Vivek Gupta 0001
WACV4
2025 Map&Make: Schema Guided Text to Table Generation
abstract
Transforming dense, unstructured text into interpretable tables-commonly referred to as Text-to-Table generation-is a key task in information extraction.Existing methods often overlook what complex information to extract and how to infer it from text.We present Map&Make, a versatile approach that decomposes text into atomic propositions to infer latent schemas, which are then used to generate tables capturing both qualitative nuances and quantitative facts.We evaluate our method on three challenging datasets: Rotowire, known for its complex, multi-table schema; Livesum which requires numerical aggregation; and Wiki40 which require open text extraction from mulitple domains.By correcting hallucination errors in Rotowire, we also provide a cleaner benchmark.Our method shows significant gains in both accuracy and interpretability across comprehensive comparative and referenceless metrics.Finally, ablation studies highlight the key factors driving performance and validate the utility of our approach in structured summarization.Code and data are available
Naman Ahuja, Fenil Denish Bardoliya, Chitta Baral, Vivek Gupta 0001
ACL (1)3
2025 GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
abstract
Publicly significant images from events carry valuable contextual information with applications in domains such as journalism and education.However, existing methodologies often struggle to accurately extract this contextual relevance from images.To address this challenge, we introduce GETREASON (Geospatial Event Temporal Reasoning), a framework designed to go beyond surfacelevel image descriptions and infer deeper contextual meaning.We hypothesize that extracting global event, temporal, and geospatial information from an image enables a more accurate understanding of its contextual significance.We also introduce a new metric GREAT (Geospatial, Reasoning and Event Accuracy with Temporal alignment) for a reasoning capturing evaluation.Our layered multi-agentic approach, evaluated using a reasoning-weighted metric, demonstrates that meaningful information can be inferred from images, allowing them to be effectively linked to their corresponding events and broader contextual background.
Shikhhar Siingh, Abhinav Rawat, Chitta Baral, Vivek Gupta 0001
ACL (1)3
2025 UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
abstract
Md Nayem Uddin, Amir Saeidi, Divij Handa, Agastya Seth, Tran Cao Son, Eduardo Blanco, Steven Corman, Chitta Baral. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Md Nayem Uddin, Amir Saeidi, Divij Handa, Agastya Seth, Tran Cao Son, Eduardo Blanco 0002, Steven R. Corman, Chitta Baral
ACL (1)8
2025 AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
abstract
Text-to-Image (T2I) models have recently achieved remarkable success in generating images from textual descriptions.However, challenges still persist in accurately rendering complex scenes where actions and interactions form the primary semantic focus.Our key observation in this work is that T2I models frequently struggle to capture nuanced and often implicit attributes inherent in action depiction, leading to generating images that lack key contextual details.To enable systematic evaluation, we introduce AcT2I, a benchmark designed to evaluate the performance of T2I models in generating images from action-centric prompts.We experimentally validate that leading T2I models do not fare well on AcT2I.We further hypothesize that this shortcoming arises from the incomplete representation of the inherent attributes and contextual dependencies in the training corpora of existing T2I models.We build upon this by developing a trainingfree, knowledge distillation technique utilizing Large Language Models to address this limitation.Specifically, we enhance prompts by incorporating dense information across three dimensions, observing that injecting prompts with temporal details significantly improves image generation accuracy, with our best model achieving an increase of 72%.Our findings highlight the limitations of current T2I methods in generating images that require complex reasoning and demonstrate that integrating linguistic knowledge in a systematic way can notably advance the generation of nuanced and contextually accurate images.
Vatsal Malaviya, Agneet Chatterjee, Maitreya Patel, Yezhou Yang, Chitta Baral
EMNLP5
2025 PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
abstract
Mihir Parmar, Palash Goyal, Xin Liu, Yiwen Song, Mingyang Ling, Chitta Baral, Hamid Palangi, Tomas Pfister. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Mihir Parmar, Palash Goyal, Yiwen Song, Chitta Baral, Hamid Palangi, Tomas Pfister
EMNLP6
2025 PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving
abstract
Mihir Parmar, Xin Liu, Palash Goyal, Yanfei Chen, Long Le, Swaroop Mishra, Hossein Mobahi, Jindong Gu, Zifeng Wang, Hootan Nakhost, Chitta Baral, Chen-Yu Lee, Tomas Pfister, Hamid Palangi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Mihir Parmar, Palash Goyal, Yanfei Chen, Long T. Le, Swaroop Mishra, Hossein Mobahi, Jindong Gu, Zifeng Wang 0002, Hootan Nakhost, Chitta Baral, Chen-Yu Lee, Tomas Pfister, Hamid Palangi
EMNLP11
2025 ThinkTuning: Instilling Cognitive Reflections without Distillation
abstract
Recent advances in test-time scaling have led to the emergence of thinking LLMs that exhibit self-reflective behaviors and multi-step reasoning.While RL drives this self-improvement paradigm, a recent study (Gandhi et al., 2025) shows that RL alone does not truly instill these new reasoning abilities -it merely draws out behaviors already present in the base models.This raises a question: How can we train models that don't exhibit such thinking behavior to develop it in the first place?To this end, we propose THINKTUNING, a GRPO-based interactive training approach where we augment the rollouts of a student model with the guidance from a teacher model.A simple idea from classroom practice inspires our method: a teacher poses a problem, lets the student try an answer, then gives corrective feedback-enough to point the mind in the right direction and then show the solution.Each piece of feedback reshapes the student's thoughts, leading them to arrive at the correct solution.Similarly, we find that this type of implicit supervision through feedback from a teacher model of the same size improves the reasoning capabilities of the student model.In particular, on average, our method shows a 3.85% improvement over zero-shot baselines across benchmarks, and on MATH-500, AIME and GPQA-Diamond it shows 2.08%, 2.23% and 3.99% improvements over the vanilla-GRPO baseline 1 .
Aswin RRV, Jacob Dineen, Divij Handa, Md Nayem Uddin, Mihir Parmar, Chitta Baral, Ben Zhou
EMNLP6
2025 RefEdit: A Benchmark and Method for Improving Instruction-Based Image Editing Model on Referring Expressions
abstract
Despite recent advances in inversion and instruction-based image editing, existing approaches primarily excel at editing single, prominent objects but significantly struggle when applied to complex scenes containing multiple entities. To quantify this gap, we first introduce RefEdit-Bench, a rigorous real-world benchmark rooted in RefCOCO, where even baselines trained on millions of samples perform poorly. To overcome this limitation, we introduce RefEdit -- an instruction-based editing model trained on our scalable synthetic data generation pipeline. Our RefEdit, trained on only 20,000 editing triplets, outperforms the Flux/SD3 model-based baselines trained on millions of data. Extensive evaluations across various benchmarks demonstrate that our model not only excels in referring expression tasks but also enhances performance on traditional benchmarks, achieving state-of-the-art results comparable to closed-source methods. We release data \& checkpoint for reproducibility.
Bimsara Pathiraja, Maitreya Patel, Shivam Singh, Yezhou Yang, Chitta Baral
ICCV5
2025 ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints
abstract
Reasoning about Actions and Change (RAC) has historically played a pivotal role in solving foundational AI problems, such as the frame problem. It has driven advancements in AI fields, such as non-monotonic and commonsense reasoning. RAC remains crucial for AI systems that operate in dynamic environments, engage in interactive scenarios, or rely on commonsense reasoning. Despite substantial advances made by Large Language Models (LLMs) in various AI domains, their performance in RAC remains underexplored. To address this gap, we introduce a new diagnostic benchmark, $\textbf{ActionReasoningBench}$, which encompasses 8 domains and includes questions for up to 19 action sequences. This benchmark rigorously evaluates LLMs across six key RAC dimensions: $\textit{Fluent Tracking}$, $\textit{State Tracking}$, $\textit{Action Executability}$, $\textit{Effects of Actions}$, $\textit{Numerical RAC}$, and $\textit{Composite Questions}$. LLMs demonstrate average accuracy rates of 73.55%, 65.63%, 58.73%, and 62.38% on the former four dimensions, which are frequently discussed in RAC literature. However, the performance on the latter two dimensions, which introduce complex and novel reasoning questions, the average performance of LLMs is lowered to 33.16% and 51.19%, respectively, reflecting a 17.9% performance decline. We also introduce new ramification constraints to capture the indirect effects of actions, providing deeper insights into RAC challenges. Our evaluation of state-of-the-art LLMs, including both open-source and commercial models, reveals challenges across all RAC dimensions, particularly in handling ramifications, with GPT-4o failing to solve any question and o1-preview achieving a score of only 18.4%.
Divij Handa, Pavel Dolin, Shrinidhi Kumbhar, Tran Cao Son, Chitta Baral
ICLR5
2025 VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
abstract
Multimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly across multiple images remains a significant challenge. To address this, we introduce VOILA, a large-scale, open-ended, dynamic benchmark designed to evaluate MLLMs' perceptual understanding and abstract relational reasoning. VOILA employs an analogical mapping approach in the visual domain, requiring models to generate an image that completes an analogy between two given image pairs, reference and application, without relying on predefined choices. Our experiments demonstrate that the analogical reasoning tasks in VOILA present a challenge to MLLMs. Through multi-step analysis, we reveal that current MLLMs struggle to comprehend inter-image relationships and exhibit limited capabilities in high-level relational reasoning. Notably, we observe that performance improves when following a multi-step strategy of least-to-most prompting. Comprehensive evaluations on open-source models and GPT-4o show that on text-based answers, the best accuracy for challenging scenarios is 13% (LLaMa 3.2) and even for simpler tasks is only 29% (GPT-4o), while human performance is significantly higher at 70% across both difficulty levels.
Nilay Yilmaz, Maitreya Patel, Yiran Luo 0001, Tejas Gokhale, Chitta Baral, Suren Jayasuriya, Yezhou Yang
ICLR5
2025 ToW: Thoughts of Words Improve Reasoning in Large Language Models
abstract
Zhikun Xu, Ming Shen, Jacob Dineen, Zhaonan Li, Xiao Ye, Shijie Lu, Aswin Rrv, Chitta Baral, Ben Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zhikun Xu, Ming Shen 0006, Jacob Dineen, Shijie Lu, Aswin RRV, Chitta Baral, Ben Zhou
NAACL (Long Papers)8
2025 Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
abstract
Recent advances in video generation have enabled high-fidelity video synthesis from user provided prompts. However, existing models and benchmarks fail to capture the complexity and requirements of professional video generation. Towards that goal, we introduce Stable Cinemetrics, a structured evaluation framework that formalizes filmmaking controls into four disentangled, hierarchical taxonomies: Setup, Event, Lighting, and Camera. Together, these taxonomies define 76 fine-grained control nodes grounded in industry practices. Using these taxonomies, we construct a benchmark of prompts aligned with professional use cases and develop an automated pipeline for prompt categorization and question generation, enabling independent evaluation of each control dimension. We conduct a large-scale human study spanning 10+ models and 20K videos, annotated by a pool of 80+ film professionals. Our analysis, both coarse and fine-grained reveal that even the strongest current models exhibit significant gaps, particularly in Events and Camera-related controls. To enable scalable evaluation, we train an automatic evaluator, a vision-language model aligned with expert annotations that outperforms existing zero-shot baselines. SCINE is the first approach to situate professional video generation within the landscape of video generative models, introducing taxonomies centered around cinematic controls and supporting them with structured evaluation pipelines and detailed analyses to guide future research.
Agneet Chatterjee, Rahim Entezari, Maksym Zhuravinskyi, Maksim Lapin, Reshinth Adithyan, Amit Raj, Chitta Baral, Yezhou Yang, Varun Jampani
NeurIPS7
2025 EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment
abstract
Erasing harmful or proprietary concepts from powerful text‑to‑image generators is an emerging safety requirement, yet current ``concept erasure'' techniques either collapse image quality, rely on brittle adversarial losses, or demand prohibitive retraining cycles. We trace these limitations to a myopic view of the denoising trajectories that govern diffusion‑based generation. We introduce EraseFlow, the first framework that casts concept unlearning as exploration in the space of denoising paths and optimizes it with a GFlowNets equipped with the trajectory‑balance objective. By sampling entire trajectories rather than single end states, EraseFlow learns a stochastic policy that steers generation away from target concepts while preserving the model’s prior. EraseFlow eliminates the need for carefully crafted reward models and by doing this, it generalizes effectively to unseen concepts and avoids hackable rewards while improving the performance. Extensive empirical results demonstrate that EraseFlow outperforms existing baselines and achieves an optimal trade-off between performance and prior preservation.
Abhiram Kusumba, Maitreya Patel, Kyle Min 0001, Changhoon Kim, Chitta Baral, Yezhou Yang
NeurIPS5
2025 Collaborative large language models for automated data extraction in living systematic reviews
abstract
OBJECTIVE: Data extraction from the published literature is the most laborious step in conducting living systematic reviews (LSRs). We aim to build a generalizable, automated data extraction workflow leveraging large language models (LLMs) that mimics the real-world 2-reviewer process. MATERIALS AND METHODS: A dataset of 10 trials (22 publications) from a published LSR was used, focusing on 23 variables related to trial, population, and outcomes data. The dataset was split into prompt development (n = 5) and held-out test sets (n = 17). GPT-4-turbo and Claude-3-Opus were used for data extraction. Responses from the 2 LLMs were considered concordant if they were the same for a given variable. The discordant responses from each LLM were provided to the other LLM for cross-critique. Accuracy, ie, the total number of correct responses divided by the total number of responses, was computed to assess performance. RESULTS: In the prompt development set, 110 (96%) responses were concordant, achieving an accuracy of 0.99 against the gold standard. In the test set, 342 (87%) responses were concordant. The accuracy of the concordant responses was 0.94. The accuracy of the discordant responses was 0.41 for GPT-4-turbo and 0.50 for Claude-3-Opus. Of the 49 discordant responses, 25 (51%) became concordant after cross-critique, increasing accuracy to 0.76. DISCUSSION: Concordant responses by the LLMs are likely to be accurate. In instances of discordant responses, cross-critique can further increase the accuracy. CONCLUSION: Large language models, when simulated in a collaborative, 2-reviewer workflow, can extract data with reasonable performance, enabling truly "living" systematic reviews.
Umair Ayub, Syed Arsalan Ahmed Naqvi, Kaneez Zahra Rubab Khakwani, Zaryab bin Riaz Sipra, Ammad Raina, Sihan Zhou, Amir Saeidi, Bashar Hasan, Robert Bryan Rumble, Danielle S. Bitterman, Jeremy L. Warner, Jia Zou 0001, Amye J. Tevaarwerk, Konstantinos Leventakos, Kenneth L. Kehl, Jeanne M. Palmer, Mohammad Hassan Murad, Chitta Baral, Irbaz Bin Riaz
J. Am. Medical Informatics Assoc.20
2024 ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
abstract
The ability to understand visual concepts and replicate and compose these concepts from images is a central goal for computer vision. Recent advances in text-to-image (T2I) models have lead to high definition and realistic image quality generation by learning from large databases of images and their descriptions. However, the evaluation of T2I models has focused on photorealism and limited qualitative measures of visual understanding. To quantify the ability of T2I models in learning and synthesizing novel visual concepts (a.k.a. personalized T2I), we introduce ConceptBed, a large-scale dataset that consists of 284 unique visual concepts, and 33K composite text prompts. Along with the dataset, we propose an evaluation metric, Concept Confidence Deviation (CCD), that uses the confidence of oracle concept classifiers to measure the alignment between concepts generated by T2I generators and concepts contained in target images. We evaluate visual concepts that are either objects, attributes, or styles, and also evaluate four dimensions of compositionality: counting, attributes, relations, and actions. Our human study shows that CCD is highly correlated with human understanding of concepts. Our results point to a trade-off between learning the concepts and preserving the compositionality which existing approaches struggle to overcome. The data, code, and interactive demo is available at: https://conceptbed.github.io/
Maitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou Yang
AAAI3
2024 LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
abstract
Mihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura, Man Luo, Santosh Mashetty, Arindam Mitra, Chitta Baral. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Mihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura, Man Luo 0003, Santosh Mashetty, Arindam Mitra, Chitta Baral
ACL (1)8
2024 On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation
abstract
Recent advances in monocular depth estimation have been made by incorporating natural language as additional guidance. Although yielding impressive results, the impact of the language prior, particularly in terms of generalization and robustness, remains unexplored. In this paper, we address this gap by quantifying the impact of this prior and introduce methods to benchmark its effectiveness across various settings. We generate “low-level” sentences that convey object-centric, three-dimensional spatial relationships, incorporate them as additional language priors and evaluate their downstream impact on depth estimation. Our key finding is that current language-guided depth estimators perform optimally only with scene-level descriptions and counter-intuitively fare worse with low level descriptions. Despite leveraging additional data, these methods are not robust to directed adversarial attacks and decline in performance with an increase in distribution shift. Finally, to provide a foundation for future research, we identify points of failures and offer insights to better understand these shortcomings. With an increasing number of methods using language for depth estimation, our findings highlight the opportunities and pitfalls that require careful consideration for effective deployment in real-world settings.11Code/Data: https://github.com/agneet42/lang_depth
Agneet Chatterjee, Tejas Gokhale, Chitta Baral, Yezhou Yang
CVPR3
2024 ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image Generations
abstract
Text-to-image (T2I) diffusion models, notably the unCLIP models (e.g., DALL-E-2), achieve state-of-the-art (SOTA) performance on various compositional T2I benchmarks, at the cost of significant computational resources. The unCLIP stack comprises T2I prior and diffusion image decoder. The T2I prior model alone adds a billion parameters compared to the Latent Diffusion Models, which increases the computational and high-quality data requirements. We introduce ECLIPSE11Our strategy, ECLIPSE, draws an analogy from the way a smaller prior model, akin to a celestial entity, offers a glimpse of the grandeur within the larger pre-trained vision-language model, mirroring how an eclipse reveals the vastness of the cosmos., a novel contrastive learning method that is both parameter and dataefficient. ECLIPSE leverages pre-trained vision-language models (e.g., CLIP) to distill the knowledge into the prior model. We demonstrate that the ECLIPSE trained prior, with only 3.3% of the parameters and trained on a mere 2.8% of the data, surpasses the baseline T2I priors with an average of 71.6% preference score under resource-limited setting. It also attains performance on par with SOTA big models, achieving an average of 63.36% preference score in terms of the ability to follow the text compositions. Extensive experiments on two unCLIP diffusion image decoders, Karlo and Kandinsky, affirm that ECLIPSE priors consistently deliver high performance while significantly reducing resource dependency. Project page: https://eclipse-t2i.vercel.app/
Maitreya Patel, Changhoon Kim, Chitta Baral, Yezhou Yang
CVPR4
2024 REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
Agneet Chatterjee, Yiran Luo 0001, Tejas Gokhale, Yezhou Yang, Chitta Baral
ECCV (30)5
2024 Getting it Right: Improving Spatial Consistency in Text-to-Image Models
Agneet Chatterjee, Gabriela Ben Melech Stan, Estelle Aflalo, Sayak Paul, Dhruba Ghosh, Tejas Gokhale, Ludwig Schmidt, Hannaneh Hajishirzi, Vasudev Lal, Chitta Baral, Yezhou Yang
ECCV (22)10
2024 Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
abstract
Nisarg Patel, Mohith Kulkarni, Mihir Parmar, Aashna Budhiraja, Mutsumi Nakamura, Neeraj Varshney, Chitta Baral. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Nisarg Patel, Mohith Kulkarni, Mihir Parmar, Aashna Budhiraja, Mutsumi Nakamura, Neeraj Varshney, Chitta Baral
EMNLP7
2024 Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
abstract
Nemika Tyagi, Mihir Parmar, Mohith Kulkarni, Aswin Rrv, Nisarg Patel, Mutsumi Nakamura, Arindam Mitra, Chitta Baral. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Nemika Tyagi, Mihir Parmar, Mohith Kulkarni, Aswin RRV, Nisarg Patel, Mutsumi Nakamura, Arindam Mitra, Chitta Baral
EMNLP8
2024 Learning Temporally Composable Task Segmentations with Language
abstract
In this work, we present an approach to identify sub-tasks within a demonstrated robot trajectory with the supervision provided by language instructions. Learning longer horizon tasks is challenging with techniques such as reinforcement learning and behavior cloning. Previous approaches have split these long tasks into shorter tasks that are easier to learn by using statistical change point detection methods. However, classical changepoint detection methods function only with low dimensional robot trajectory data and not with high dimensional inputs such as vision. Our goal in this work is to split longer horizon tasks, represented by trajectories into shorter horizon tasks that can be learned using conventional behavior cloning approaches using guidance from language. In our approach we use techniques from the video moment retrieval problem on robot trajectory data to demonstrate a high-dimensional generalizable change-point detection approach. Our proposed moment retrieval-based approach shows a more than 30% improvement in mean average precision (mAP) for identifying trajectory sub-tasks with language guidance compared to that without language. We perform ablations to understand the effects of domain randomization, sample complexity, views, and sim-to-real transfer of our method. In our data ablation we find that just with a 100 labelled trajectories we can achieve a 61.41 mAP, demonstrating the sample efficiency of using such an approach. Further, behavior cloning models trained on our segmented trajectories outperform a single model trained on the whole trajectory by up to 20%.
Divyanshu Raj, Omkar Patil, Weiwei Gu, Chitta Baral, Nakul Gopalan
IROS4
2024 TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
abstract
Contrastive Language-Image Pretraining (CLIP) models maximize the mutual information between text and visual modalities to learn representations. This makes the nature of the training data a significant factor in the efficacy of CLIP for downstream tasks. However, the lack of compositional diversity in contemporary image-text datasets limits the compositional reasoning ability of CLIP. We show that generating ``hard'' negative captions via in-context learning and synthesizing corresponding negative images with text-to-image generators offers a solution. We introduce a novel contrastive pre-training strategy that leverages these hard negative captions and images in an alternating fashion to train CLIP. We demonstrate that our method, named TripletCLIP, when applied to existing datasets such as CC3M and CC12M, enhances the compositional capabilities of CLIP, resulting in an absolute improvement of over 9% on the SugarCrepe benchmark on an equal computational budget, as well as improvements in zero-shot image classification and image retrieval. Our code, models, and data are available at: tripletclip.github.io.
Maitreya Patel, Abhiram Kusumba, Changhoon Kim, Tejas Gokhale, Chitta Baral, Yezhou Yang
NeurIPS6
2024 "Len or index or count, anything but v1": Predicting Variable Names in Decompilation Output with Transfer Learning
abstract
Binary reverse engineering is an arduous and tedious task performed by skilled and expensive human analysts. Information about the source code is irrevocably lost in the compilation process. While modern decompilers attempt to generate C-style source code from a binary, they cannot recover lost variable names. Prior works have explored machine learning techniques for predicting variable names in decompiled code. However, the state-of-the-art systems, DIRE and DIRTY, generalize poorly to functions in the testing set that are not included in the training set—31.8% for DIRE on DIRTY’s data set and 36.9% for DIRTY on DIRTY’s data set.In this paper, we present VarBERT, a Bidirectional Encoder Representations from Transformers (BERT) to predict meaningful variable names in decompilation output. An advantage of VarBERT is that we can pre-train on human source code and then fine-tune the model to the task of predicting variable names. We also create a new data set VarCorpus, which significantly expands the size and variety of the data set. Our evaluation of VarBERT on VarCorpus, demonstrates a significant improvement in predicting the developer’s original variable names for O2 optimized binaries achieving accuracies of 54.43% for IDA and 54.49% for Ghidra. VarBERT is strictly better than state-of-the-art techniques: On a subset of VarCorpus, VarBERT could predict the developer’s original variable names 50.70% of the time, while DIRE and DIRTY predicted original variable names 35.94% and 38.00% of the time, respectively.
Kuntal Kumar Pal, Ati Priya Bajaj, Pratyay Banerjee, Audrey Dutcher, Mutsumi Nakamura, Zion Leonahenahe Basque, Saurabh Arjun Sawant, Ujjwala Anantheswaran, Yan Shoshitaishvili, Adam Doupé, Chitta Baral, Ruoyu Wang 0001
SP12
2023 End-to-end Knowledge Retrieval with Multi-modal Queries
abstract
We investigate knowledge retrieval with multimodal queries, i.e. queries containing information split across image and text inputs, a challenging task that differs from previous work on cross-modal retrieval.We curate a new dataset called ReMuQ 1 for benchmarking progress on this task.ReMuQ requires a system to retrieve knowledge from a large corpus by integrating contents from both text and image queries.We introduce a retriever model "ReViz" that can directly process input text and images to retrieve relevant knowledge in an end-to-end fashion without being dependent on intermediate modules such as object detectors or caption generators.We introduce a new pretraining task that is effective for learning knowledge retrieval with multimodal queries and also improves performance on downstream tasks.We demonstrate superior performance in retrieval on two datasets (ReMuQ and OK-VQA) under zeroshot settings as well as further improvements when finetuned on these datasets.
Man Luo 0003, Zhiyuan Fang, Tejas Gokhale, Yezhou Yang, Chitta Baral
ACL (1)5
2023 Post-Abstention: Towards Reliably Re-Attempting the Abstained Instances in QA
abstract
Despite remarkable progress made in natural language processing, even the state-of-the-art models often make incorrect predictions.Such predictions hamper the reliability of systems and limit their widespread adoption in realworld applications.Selective prediction partly addresses the above concern by enabling models to abstain from answering when their predictions are likely to be incorrect.While selective prediction is advantageous, it leaves us with a pertinent question 'what to do after abstention'.To this end, we present an explorative study on 'Post-Abstention', a task that allows re-attempting the abstained instances with the aim of increasing coverage of the system without significantly sacrificing its accuracy.We first provide mathematical formulation of this task and then explore several methods to solve it.Comprehensive experiments on 11 QA datasets show that these methods lead to considerable risk improvements -performance metric of the Post-Abstention task-both in the in-domain and the out-of-domain settings.We also conduct a thorough analysis of these results which further leads to several interesting findings.Finally, we believe that our work will encourage and facilitate further research in this important area of addressing the reliability of NLP systems.
Neeraj Varshney, Chitta Baral
ACL (1)2
2023 Real-Time Visual Feedback to Guide Benchmark Creation: A Human-and-Metric-in-the-Loop Workflow
abstract
Recent research has shown that language models exploit 'artifacts' in benchmarks to solve tasks, rather than truly learning them, leading to inflated model performance.In pursuit of creating better benchmarks, we propose VAIDA, a novel benchmark creation paradigm for NLP, that focuses on guiding crowdworkers, an under-explored facet of addressing benchmark idiosyncrasies.VAIDA facilitates sample correction by providing real-time visual feedback and recommendations to improve sample quality.Our approach is domain, model, task, and metric agnostic, and constitutes a paradigm shift for robust, validated, and dynamic benchmark creation via human-and-metric-in-theloop workflows.We evaluate via expert review and a user study with NASA TLX.We find that VAIDA decreases effort, frustration, mental, and temporal demands of crowdworkers and analysts, simultaneously increasing the performance of both user groups with a 45.8% decrease in the level of artifacts in created samples.As a by-product of our user study, we observe that created samples are adversarial across models, leading to decreases of 31.3% (BERT), 22.5% (RoBERTa), 14.98% (GPT-3 fewshot) in performance.1
Anjana Arunkumar, Swaroop Mishra, Bhavdeep Singh Sachdeva, Chitta Baral, Chris Bryan
EACL4
2023 "John is 50 years old, can his son be 65?" Evaluating NLP Models' Understanding of Feasibility
abstract
Himanshu Gupta, Neeraj Varshney, Swaroop Mishra, Kuntal Kumar Pal, Saurabh Arjun Sawant, Kevin Scaria, Siddharth Goyal, Chitta Baral. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Neeraj Varshney, Swaroop Mishra, Kuntal Kumar Pal, Saurabh Arjun Sawant, Kevin Scaria, Siddharth Goyal, Chitta Baral
EACL8
2023 Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
abstract
In recent years, progress in NLU has been driven by benchmarks.These benchmarks are typically collected by crowdsourcing, where annotators write examples based on annotation instructions crafted by dataset creators.In this work, we hypothesize that annotators pick up on patterns in the crowdsourcing instructions, which bias them to write many similar examples that are then over-represented in the collected data.We study this form of bias, termed instruction bias, in 14 recent NLU benchmarks, showing that instruction examples often exhibit concrete patterns, which are propagated by crowdworkers to the collected data.This extends previous work (Geva et al., 2019) and raises a new concern of whether we are modeling the dataset creator's instructions, rather than the task.Through a series of experiments, we show that, indeed, instruction bias can lead to overestimation of model performance, and that models struggle to generalize beyond biases originating in the crowdsourcing instructions.We further analyze the influence of instruction bias in terms of pattern frequency and model size, and derive concrete recommendations for creating future NLU benchmarks. 1
Mihir Parmar, Swaroop Mishra, Mor Geva, Chitta Baral
EACL4
2023 Improving Diversity with Adversarially Learned Transformations for Domain Generalization
abstract
To be successful in single source domain generalization (SSDG), maximizing diversity of synthesized domains has emerged as one of the most effective strategies. Recent success in SSDG comes from methods that pre-specify diversity inducing image augmentations during training, so that it may lead to better generalization on new domains. However, naïve pre-specified augmentations are not always effective, either because they cannot model large domain shift, or be-cause the specific choice of transforms may not cover the types of shift commonly occurring in domain generalization. To address this issue, we present a novel framework called ALT: adversarially learned transformations, that uses an adversary neural network to model plausible, yet hard image transformations that fool the classifier. ALT learns image transformations by randomly initializing the adversary net-work for each batch and optimizing it for a fixed number of steps to maximize classification error. The classifier is trained by enforcing a consistency between its predictions on the clean and transformed images. With extensive empirical analysis, we find that this new form of adversarial transformations achieves both objectives of diversity and hardness simultaneously, outperforming all existing techniques on competitive benchmarks for SSDG. We also show that ALT can seamlessly work with existing diversity modules to produce highly distinct, and large transformations of the source domain leading to state-of-the-art performance. Code: https://github.com/tejas-gokhale/ALT
Tejas Gokhale, Rushil Anirudh, Jayaraman J. Thiagarajan, Bhavya Kailkhura, Chitta Baral, Yezhou Yang
WACV5
2022 Improving Biomedical Information Retrieval with Neural Retrievers
abstract
Information retrieval (IR) is essential in search engines and dialogue systems as well as natural language processing tasks such as open-domain question answering. IR serve an important function in the biomedical domain, where content and sources of scientific knowledge may evolve rapidly. Although neural retrievers have surpassed traditional IR approaches such as TF-IDF and BM25 in standard open-domain question answering tasks, they are still found lacking in the biomedical domain. In this paper, we seek to improve information retrieval (IR) using neural retrievers (NR) in the biomedical domain, and achieve this goal using a three-pronged approach. First, to tackle the relative lack of data in the biomedical domain, we propose a template-based question generation method that can be leveraged to train neural retriever models. Second, we develop two novel pre-training tasks that are closely aligned to the downstream task of information retrieval. Third, we introduce the ``Poly-DPR'' model which encodes each context into multiple context vectors. Extensive experiments and analysis on the BioASQ challenge suggest that our proposed method leads to large gains over existing neural approaches and beats BM25 in the small-corpus setting. We show that BM25 and our method can complement each other, and a simple hybrid model leads to further gains in the large corpus setting.
Man Luo 0003, Arindam Mitra, Tejas Gokhale, Chitta Baral
AAAI4
2022 Cross-Task Generalization via Natural Language Crowdsourcing Instructions
abstract
Humans (e.g., crowdworkers) have a remarkable ability in solving different tasks, by simply reading textual instructions that define them and looking at a few examples.Despite the success of the conventional supervised learning on individual datasets, such models often struggle with generalization across tasks (e.g., a question-answering system cannot solve classification tasks).A long-standing challenge in AI is to build a model that learns a new task by understanding the humanreadable instructions that define it.To study this, we introduce NATURAL INSTRUCTIONS, a dataset of 61 distinct tasks, their humanauthored instructions, and 193k task instances (input-output pairs).The instructions are obtained from crowdsourcing instructions used to create existing NLP datasets and mapped to a unified schema.Using this meta-dataset, we measure cross-task generalization by training models on seen tasks and measuring generalization to the remaining unseen ones.We adopt generative pre-trained language models to encode task-specific instructions along with input and generate task output.Our results indicate that models benefit from instructions when evaluated in terms of generalization to unseen tasks (19% better for models utilizing instructions).These models, however, are far behind an estimated performance upperbound, indicating significant room for more progress in this direction.1
Swaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh Hajishirzi
ACL (1)3
2022 NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks
abstract
Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Sachdeva, Peter Clark, Chitta Baral, Ashwin Kalyan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Singh Sachdeva, Peter Clark, Chitta Baral, Ashwin Kalyan
ACL (1)6
2022 ILDAE: Instance-Level Difficulty Analysis of Evaluation Data
abstract
Knowledge of difficulty level of questions helps a teacher in several ways, such as estimating students' potential quickly by asking carefully selected questions and improving quality of examination by modifying trivial and hard questions.Can we extract such benefits of instance difficulty in Natural Language Processing?To this end, we conduct Instance-Level Difficulty Analysis of Evaluation data (ILDAE) in a largescale setup of 23 datasets and demonstrate its five novel applications: 1) conducting efficientyet-accurate evaluations with fewer instances saving computational cost and time, 2) improving quality of existing evaluation datasets by repairing erroneous and trivial instances, 3) selecting the best model based on application requirements, 4) analyzing dataset characteristics for guiding future data creation, 5) estimating Out-of-Domain performance reliably.Comprehensive experiments for these applications lead to several interesting results, such as evaluation using just 5% instances (selected via ILDAE) achieves as high as 0.93 Kendall correlation with evaluation using complete dataset and computing weighted accuracy using difficulty scores leads to 5.2% higher correlation with Out-of-Domain performance.We release the difficulty scores 1 and hope our work will encourage research in this important yet understudied field of leveraging instance difficulty in evaluations.
Neeraj Varshney, Swaroop Mishra, Chitta Baral
ACL (1)3
2022 Less is More: Summary of Long Instructions is Better for Program Synthesis
abstract
Despite the success of large pre-trained language models (LMs) such as Codex, they show below-par performance on the larger and more complicated programming related questions.We show that LMs benefit from the summarized version of complicated questions.Our findings show that superfluous information often present in problem description such as human characters, background stories, and names (which are included to help humans in understanding a task) does not help models in understanding a task.To this extent, we create a metadataset from the frequently used APPS dataset and the newly created CodeContests dataset for the program synthesis task.Our meta-dataset consists of human and synthesized summaries of the long and complicated programming questions.Experimental results on Codex show that our proposed approach outperforms baseline by 8.13% on the APPS dataset and 11.88% on the CodeContests dataset on average in terms of strict accuracy.Our analysis shows that summaries significantly improve performance for introductory (9.86%) and interview (11.48%) programming questions.However, it shows improvement by a small margin (∼ 2%) for competitive programming questions, implying scope for future research in this direction.1 * Equal Contribution 1 Code and data is available at https://github.com/kurbster/ Prompt-Summarization 2 Detailed related work is presented in Appendix A However, LMs such as Codex show below-par performance on the long and complicated programming questions.We observe that the natural language description of the program becomes long and complicated when there is superfluous information (see section 2.1.1).The goal of adding this information to the description is to make it more understandable to humans.However, we find that this information confuses the model in understanding a task 3 .We propose that removing the excess information and providing the model with the exact specifications of the problem can improve the performance of the LMs.To remove excess information 4 , we summarize the descriptions of the program in such a way that it does not lose important specifications.We use the APPS dataset (Hendrycks et al., 2021) and Code-Contests dataset (Li et al., 2022) which are a collection of coding problems from different online sources and create a meta-dataset consisting of human and synthesized summaries.We perform all experiments using the GPT-based Codex model (Chen et al., 2021) on the proposed meta-dataset and show that the summarized version of complicated questions improves strict accuracy by 8.13% on the APPS dataset and 11.85% on CodeContests.From our analysis, we can see significant improvement for introductory (9.86%) and interview (11.48%) related programming questions.However, it shows improvement by a small margin (∼ 2%) for competitive programming questions.Considering that automatic evaluation of a program does not reward for partial correctness, we perform qualitative evaluation on our meta-dataset and find that original questions often confuse models in understanding the underlying problem, as models latch on to some spurious words in the text (e.g. the word 'list' in question makes the model 3 See example in Appendix C 4 Instructions for creating summaries given in Appendix N 4532 design a list even though the underlying problem is on graphs).We further analyze model performance on different types of summaries (i.e., basic, expert, and synthetic) and provide instruction-design principles that can help future research on prompting in program synthesis.2 Method 2.1 Dataset We use the APPS (Hendrycks et al., 2021) and CodeContests (Li et al., 2022) datasets to create summaries.We crowd-sourced the creation of human summaries.The result was 373 human summaries for APPS and 80 summaries for CodeContests along with and 8663 synthetic summaries using both datasets.Table 1 shows the statistics of the generated summaries.
Kirby Kuznia, Swaroop Mishra, Mihir Parmar, Chitta Baral
EMNLP4
2022 LILA: A Unified Benchmark for Mathematical Reasoning
abstract
Swaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, Ashwin Kalyan. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Swaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, Ashwin Kalyan
EMNLP6
2022 CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering
abstract
Videos often capture objects, their visible properties, their motion, and the interactions between different objects.Objects also have physical properties such as mass, which the imaging pipeline is unable to directly capture.However, these properties can be estimated by utilizing cues from relative object motion and the dynamics introduced by collisions.In this paper, we introduce CRIPP-VQA 1 , a new video question answering dataset for reasoning about the implicit physical properties of objects in a scene.CRIPP-VQA contains videos of objects in motion, annotated with questions that involve counterfactual reasoning about the effect of actions, questions about planning in order to reach a goal, and descriptive questions about visible properties of objects.The CRIPP-VQA test set enables evaluation under several outof-distribution settings -videos with objects with masses, coefficients of friction, and initial velocities that are not observed in the training distribution.Our experiments reveal a surprising and significant performance gap in terms of answering questions about implicit properties (the focus of this paper) and explicit properties of objects (the focus of prior work).
Maitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou Yang
EMNLP3
2022 Is a Question Decomposition Unit All We Need?
abstract
Large Language Models (LMs) have achieved state-of-the-art performance on many Natural Language Processing (NLP) benchmarks.With the growing number of new benchmarks, we build bigger and more complex LMs.However, building new LMs may not be an ideal option owing to the cost, time and environmental impact associated with it.We explore an alternative route: can we modify data by expressing it in terms of the model's strengths, so that a question becomes easier for models to answer?We investigate if humans can decompose a hard question into a set of simpler questions that are relatively easier for models to solve.We analyze a range of datasets involving various forms of reasoning and find that it is indeed possible to significantly improve model performance (24% for GPT3 and 29% for RoBERTa-SQuAD along with a symbolic calculator) via decomposition.Our approach provides a viable option to involve people in NLP research in a meaningful way.Our findings indicate that Human-in-the-loop Question Decomposition (HQD) can potentially provide an alternate path to building large LMs 1 .
Pruthvi Patel, Swaroop Mishra, Mihir Parmar, Chitta Baral
EMNLP4
2022 Model Cascading: Towards Jointly Improving Efficiency and Accuracy of NLP Systems
abstract
Do all instances need inference through the big models for a correct prediction?Perhaps not; some instances are easy and can be answered correctly by even small capacity models.This provides opportunities for improving the computational efficiency of systems.In this work, we present an explorative study on 'model cascading', a simple technique that utilizes a collection of models of varying capacities to accurately yet efficiently output predictions.Through comprehensive experiments in multiple task settings that differ in the number of models available for cascading (K value), we show that cascading improves both the computational efficiency and the prediction accuracy.For instance, in K=3 setting, cascading saves up to 88.93% computation cost and consistently achieves superior prediction accuracy with an improvement of up to 2.18%.We also study the impact of introducing additional models in the cascade and show that it further increases the efficiency improvements.Finally, we hope that our work will facilitate development of efficient NLP systems making their widespread adoption in real-world applications possible.
Neeraj Varshney, Chitta Baral
EMNLP2
2022 An action language for multi-agent domains
Chitta Baral, Gregory Gelfond, Enrico Pontelli, Tran Cao Son
Artif. Intell.1
2021 Attribute-Guided Adversarial Training for Robustness to Natural Perturbations
abstract
While existing work in robust deep learning has focused on small pixel-level norm-based perturbations, this may not account for perturbations encountered in several real world settings. In many such cases although test data might not be available, broad specifications about the types of perturbations (such as an unknown degree of rotation) may be known. We consider a setup where robustness is expected over an unseen test domain that is not i.i.d. but deviates from the training domain. While this deviation may not be exactly known, its broad characterization is specified a priori, in terms of attributes. We propose an adversarial training approach which learns to generate new samples so as to maximize exposure of the classifier to the attributes-space, without having access to the data from the test domain. Our adversarial training solves a min-max optimization problem, with the inner maximization generating adversarial perturbations, and the outer minimization finding model parameters by optimizing the loss on adversarial perturbations generated from the inner maximization. We demonstrate the applicability of our approach on three types of naturally occurring perturbations --- object-related shifts, geometric transformations, and common image corruptions. Our approach enables deep neural networks to be robust against a wide range of naturally occurring perturbations. We demonstrate the usefulness of the proposed approach by showing the robustness gains of deep neural networks trained using our adversarial training on MNIST, CIFAR-10, and a new variant of the CLEVR dataset.
Tejas Gokhale, Rushil Anirudh, Bhavya Kailkhura, Jayaraman J. Thiagarajan, Chitta Baral, Yezhou Yang
AAAI5
2021 'Just because you are right, doesn't mean I am wrong': Overcoming a bottleneck in development and evaluation of Open-Ended VQA tasks
abstract
Man Luo, Shailaja Keyur Sampat, Riley Tallman, Yankai Zeng, Manuha Vancha, Akarshan Sajja, Chitta Baral. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Man Luo 0003, Shailaja Sampat, Riley Tallman, Yankai Zeng, Manuha Vancha, Akarshan Sajja, Chitta Baral
EACL7
2021 Weakly-Supervised Visual-Retriever-Reader for Knowledge-based Question Answering
abstract
Knowledge-based visual question answering (VQA) requires answering questions with external knowledge in addition to the content of images.One dataset that is mostly used in evaluating knowledge-based VQA is OK-VQA, but it lacks a gold standard knowledge corpus for retrieval.Existing work leverage different knowledge bases (e.g., ConceptNet and Wikipedia) to obtain external knowledge.Because of varying knowledge bases, it is hard to fairly compare models' performance.To address this issue, we collect a natural language knowledge base that can be used for any VQA system.Moreover, we propose a Visual Retriever-Reader pipeline to approach knowledge-based VQA.The visual retriever aims to retrieve relevant knowledge, and the visual reader seeks to predict answers based on given knowledge.We introduce various ways to retrieve knowledge using text and images and two reader styles: classification and extraction.Both the retriever and reader are trained with weak supervision.Our experimental results show that a good retriever can significantly improve the reader's performance on the OK-VQA challenge.The code and corpus are provided in this link.
Man Luo 0003, Yankai Zeng, Pratyay Banerjee, Chitta Baral
EMNLP (1)4
2021 Weakly Supervised Relative Spatial Reasoning for Visual Question Answering
abstract
Vision-and-language (V&L) reasoning necessitates perception of visual concepts such as objects and actions, understanding semantics and language grounding, and reasoning about the interplay between the two modalities. One crucial aspect of visual reasoning is spatial understanding, which involves understanding relative locations of objects, i.e. implicitly learning the geometry of the scene. In this work, we evaluate the faithfulness of V&L models to such geometric understanding, by formulating the prediction of pair-wise relative locations of objects as a classification as well as a regression task. Our findings suggest that state-of-the-art transformer-based V&L models lack sufficient abilities to excel at this task. Motivated by this, we design two objectives as proxies for 3D spatial reasoning (SR) – object centroid estimation, and relative position estimation, and train V&L with weak supervision from off-the-shelf depth estimators. This leads to considerable improvements in accuracy for the "GQA" visual question answering challenge (in fully supervised, few-shot, and O.O.D settings) as well as improvements in relative spatial reasoning. Code and data will be released here.
Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, Chitta Baral
ICCV4
2021 Self-Supervised Test-Time Learning for Reading Comprehension
abstract
Recent work on unsupervised question answering has shown that models can be trained with procedurally generated question-answer pairs and can achieve performance competitive with supervised methods.In this work, we consider the task of unsupervised reading comprehension and present a method that performs "test-time learning" (TTL) on a given context (text passage), without requiring training on large-scale human-authored datasets containing context-question-answer triplets.This method operates directly on a single test context, uses self-supervision to train models on synthetically generated question-answer pairs, and then infers answers to unseen humanauthored questions for this context.Our method achieves accuracies competitive with fully supervised methods and significantly outperforms current unsupervised methods.TTL methods with a smaller model are also competitive with the current state-of-the-art in unsupervised reading comprehension.
Pratyay Banerjee, Tejas Gokhale, Chitta Baral
NAACL-HLT3
2021 CLEVR_HYP: A Challenge Dataset and Baselines for Visual Question Answering with Hypothetical Actions over Images
abstract
Shailaja Keyur Sampat, Akshay Kumar, Yezhou Yang, Chitta Baral. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Shailaja Sampat, Yezhou Yang, Chitta Baral
NAACL-HLT4
2021 Biomedical Named Entity Recognition via Knowledge Guidance and Question Answering
abstract
In this work, we formulated the named entity recognition (NER) task as a multi-answer knowledge guided question-answer task (KGQA) and showed that the knowledge guidance helps to achieve state-of-the-art results for 11 of 18 biomedical NER datasets. We prepended five different knowledge contexts—entity types, questions, definitions, and examples—to the input text and trained and tested BERT-based neural models on such input sequences from a combined dataset of the 18 different datasets. This novel formulation of the task (a) improved named entity recognition and illustrated the impact of different knowledge contexts, (b) reduced system confusion by limiting prediction to a single entity-class for each input token (i.e.,B,I,Oonly) compared to multiple entity-classes in traditional NER (i.e.,Bentity1,Bentity2,Ientity1,I,O), (c) made detection of nested entities easier, and (d) enabled the models to jointly learn NER-specific features from a large number of datasets. We performed extensive experiments of this KGQA formulation on the biomedical datasets, and through the experiments, we showed when knowledge improved named entity recognition. We analyzed the effect of the task formulation, the impact of the different knowledge contexts, the multi-task aspect of the generic format, and the generalization ability of KGQA. We also probed the model to better understand the key contributors for these improvements.
Pratyay Banerjee, Kuntal Kumar Pal, Murthy V. Devarakonda, Chitta Baral
ACM Trans. Comput. Heal.4
2020 Enhancing Natural Language Inference Using New and Expanded Training Data Sets and New Learning Models
abstract
Natural Language Inference (NLI) plays an important role in many natural language processing tasks such as question answering. However, existing NLI modules that are trained on existing NLI datasets have several drawbacks. For example, they do not capture the notion of entity and role well and often end up making mistakes such as “Peter signed a deal” can be inferred from “John signed a deal”. As part of this work, we have developed two datasets that help mitigate such issues and make the systems better at understanding the notion of “entities” and “roles”. After training the existing models on the new dataset we observe that the existing models do not perform well on one of the new benchmark. We then propose a modification to the “word-to-word” attention function which has been uniformly reused across several popular NLI architectures. The resulting models perform as well as their unmodified counterparts on the existing benchmarks and perform significantly well on the new benchmarks that emphasize “roles” and “entities”.
Arindam Mitra, Ishan Shrivastava, Chitta Baral
AAAI3
2020 VQA-LOL: Visual Question Answering Under the Lens of Logic
Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang
ECCV (21)3
2020 Self-Supervised Knowledge Triplet Learning for Zero-Shot Question Answering
abstract
The aim of all Question Answering (QA) systems is to generalize to unseen questions.Current supervised methods are reliant on expensive data annotation.Moreover, such annotations can introduce unintended annotator bias, making systems focus more on the bias than the actual task.This work proposes Knowledge Triplet Learning (KTL), a self-supervised task over knowledge graphs.We propose heuristics to create synthetic graphs for commonsense and scientific knowledge.We propose using KTL to perform zero-shot question answering, and our experiments show considerable improvements over large pre-trained transformer language models.
Pratyay Banerjee, Chitta Baral
EMNLP (1)2
2020 Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning
abstract
Captioning is a crucial and challenging task for video understanding.In videos that involve active agents such as humans, the agent's actions can bring about myriad changes in the scene.Observable changes such as movements, manipulations, and transformations of the objects in the scene, are reflected in conventional video captioning.Unlike images, actions in videos are also inherently linked to social aspects such as intentions (why the action is taking place), effects (what changes due to the action), and attributes that describe the agent.Thus for video understanding, such as when captioning videos or when answering questions about videos, one must have an understanding of these commonsense aspects.We present the first work on generating commonsense captions directly from videos, to describe latent aspects such as intentions, effects, and attributes.We present a new dataset "Video-to-Commonsense (V2C)" that contains ∼ 9k videos of human agents performing various actions, annotated with 3 types of commonsense descriptions.Additionally we explore the use of open-ended video-based commonsense question answering (V2C-QA) as a way to enrich our captions.Both the generation task and the QA task can be used to enrich video captions.. frame frame frame CNN
Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang
EMNLP (1)4
2020 MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question Answering
abstract
While progress has been made on the visual question answering leaderboards, models often utilize spurious correlations and priors in datasets under the i.i.d.setting.As such, evaluation on out-of-distribution (OOD) test samples has emerged as a proxy for generalization.In this paper, we present MUTANT, a training paradigm that exposes the model to perceptually similar, yet semantically distinct mutations of the input, to improve OOD generalization, such as the VQA-CP challenge.Under this paradigm, models utilize a consistency-constrained training objective to understand the effect of semantic changes in input (question-image pair) on the output (answer).Unlike existing methods on VQA-CP, MUTANT does not rely on the knowledge about the nature of train and test answer distributions.MUTANT establishes a new state-ofthe-art accuracy on VQA-CP with a 10.57% improvement.Our work opens up avenues for the use of semantic input mutations for OOD generalization in question answering.
Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang
EMNLP (1)3
2020 Language-Conditioned Imitation Learning for Robot Manipulation Tasks
abstract
Imitation learning is a popular approach for teaching motor skills to robots. However, most approaches focus on extracting policy parameters from execution traces alone (i.e., motion trajectories and perceptual data). No adequate communication channel exists between the human expert and the robot to describe critical aspects of the task, such as the properties of the target object or the intended shape of the motion. Motivated by insights into the human teaching process, we introduce a method for incorporating unstructured natural language into imitation learning. At training time, the expert can provide demonstrations along with verbal descriptions in order to describe the underlying intent (e.g., "go to the large green bowl"). The training process then interrelates these two modalities to encode the correlations between language, perception, and motion. The resulting language-conditioned visuomotor policies can be conditioned at runtime on new human commands and instructions, which allows for more fine-grained control over the trained policies while also reducing situational ambiguity. We demonstrate in a set of simulation experiments how our approach can learn language-conditioned manipulation policies for a seven-degree-of-freedom robot arm and compare the results to a variety of alternative methods.
Simon Stepputtis, Joseph Campbell, Mariano J. Phielipp, Stefan Lee, Chitta Baral, Heni Ben Amor
NeurIPS5
2019 Declarative Question Answering over Knowledge Bases Containing Natural Language Text with Answer Set Programming
abstract
While in recent years machine learning (ML) based approaches have been the popular approach in developing endto-end question answering systems, such systems often struggle when additional knowledge is needed to correctly answer the questions. Proposed alternatives involve translating the question and the natural language text to a logical representation and then use logical reasoning. However, this alternative falters when the size of the text gets bigger. To address this we propose an approach that does logical reasoning over premises written in natural language text. The proposed method uses recent features of Answer Set Programming (ASP) to call external NLP modules (which may be based on ML) which perform simple textual entailment. To test our approach we develop a corpus based on the life cycle questions and showed that Our system achieves up to 18% performance gain when compared to standard MCQ solvers.
Arindam Mitra, Peter Clark, Oyvind Tafjord, Chitta Baral
AAAI4
2019 Careful Selection of Knowledge to Solve Open Book Question Answering
abstract
Open book question answering is a type of natural language based QA (NLQA) where questions are expected to be answered with respect to a given set of open book facts, and common knowledge about a topic. Recently a challenge involving such QA, OpenBookQA, has been proposed. Unlike most other NLQA that focus on linguistic understanding, OpenBookQA requires deeper reasoning involving linguistic understanding as well as reasoning with common knowledge. In this paper we address QA with respect to the OpenBookQA dataset and combine state of the art language models with abductive information retrieval (IR), information gain based re-ranking, passage selection and weighted scoring to achieve 72.0% accuracy, an 11.6% improvement over the current state of the art.
Pratyay Banerjee, Kuntal Kumar Pal, Arindam Mitra, Chitta Baral
ACL (1)4
2019 Combining Knowledge Hunting and Neural Language Models to Solve the Winograd Schema Challenge
abstract
Winograd Schema Challenge (WSC) is a pronoun resolution task which seems to require reasoning with commonsense knowledge.The needed knowledge is not present in the given text.Automatic extraction of the needed knowledge is a bottleneck in solving the challenge.The existing state-of-the-art approach uses the knowledge embedded in their pretrained language model.However, the language models only embed part of the knowledge, the ones related to frequently co-existing concepts.This limits the performance of such models on the WSC problems.In this work, we build-up on the language model based methods and augment them with a commonsense knowledge hunting (using automatic extraction from text) module and an explicit reasoning module.Our end-to-end system built in such a manner improves on the accuracy of two of the available language model based approaches by 5.53% and 7.7% respectively.Overall our system achieves the state-of-theart accuracy of 71.06% on the WSC dataset, an improvement of 7.36% over the previous best.
Ashok Prakash, Arpit Sharma 0001, Arindam Mitra, Chitta Baral
ACL (1)4
2019 Integrating Knowledge and Reasoning in Image Understanding
abstract
Deep learning based data-driven approaches have been successfully applied in various image understanding applications ranging from object recognition, semantic segmentation to visual question answering. However, the lack of knowledge integration as well as higher-level reasoning capabilities with the methods still pose a hindrance. In this work, we present a brief survey of a few representative reasoning mechanisms, knowledge integration methods and their corresponding image understanding applications developed by various groups of researchers, approaching the problem from a variety of angles. Furthermore, we discuss upon key efforts on integrating external knowledge with neural networks. Taking cues from these efforts, we conclude by discussing potential pathways to improve reasoning capabilities.
Somak Aditya, Yezhou Yang, Chitta Baral
IJCAI3
2019 Spatial Knowledge Distillation to Aid Visual Reasoning
abstract
For tasks involving language and vision, the current state-of-the-art methods tend not to leverage any additional information that might be present to gather relevant (commonsense) knowledge. A representative task is Visual Question Answering where large diagnostic datasets have been proposed to test a system's capability of answering questions about images. The training data is often accompanied by annotations of individual object properties and spatial locations. In this work, we take a step towards integrating this additional privileged information in the form of spatial knowledge to aid in visual reasoning. We propose a framework that combines recent advances in knowledge distillation (teacher-student framework), relational reasoning and probabilistic logical languages to incorporate such knowledge in existing neural networks for the task of Visual Question Answering. Specifically, for a question posed against an image, we use a probabilistic logical language to encode the spatial knowledge and the spatial understanding about the question in the form of a mask that is directly provided to the teacher network. The student network learns from the ground-truth information as well as the teachers prediction via distillation. We also demonstrate the impact of predicting such a mask inside the teachers network using attention. Empirically, we show that both the methods improve the test accuracy over a state-of-the-art approach on a publicly available dataset.
Somak Aditya, Rudra Saha, Yezhou Yang, Chitta Baral
WACV4
2018 Explicit Reasoning over End-to-End Neural Architectures for Visual Question Answering
abstract
Many vision and language tasks require commonsense reasoning beyond data-driven image and natural language processing. Here we adopt Visual Question Answering (VQA) as an example task, where a system is expected to answer a question in natural language about an image. Current state-of-the-art systems attempted to solve the task using deep neural architectures and achieved promising performance. However, the resulting systems are generally opaque and they struggle in understanding questions for which extra knowledge is required. In this paper, we present an explicit reasoning layer on top of a set of penultimate neural network based systems. The reasoning layer enables reasoning and answering questions where additional knowledge is required, and at the same time provides an interpretable interface to the end users. Specifically, the reasoning layer adopts a Probabilistic Soft Logic (PSL) based engine to reason over a basket of inputs: visual relations, the semantic parse of the question, and background ontological knowledge from word2vec and ConceptNet. Experimental analysis of the answers and the key evidential predicates generated on the VQA dataset validate our approach.
Somak Aditya, Yezhou Yang, Chitta Baral
AAAI3
2018 Knowledge Representation and Reasoning in Answering Science Questions: A Case Study for Food Web Questions
Arindam Mitra, Chitta Baral, Peter Clark
KR2
2018 Combining Knowledge and Reasoning through Probabilistic Soft Logic for Image Puzzle Solving
Somak Aditya, Yezhou Yang, Chitta Baral, Yiannis Aloimonos
UAI3
2018 Image Understanding using vision and reasoning through Scene Description Graph
Somak Aditya, Yezhou Yang, Chitta Baral, Yiannis Aloimonos, Cornelia Fermüller
Comput. Vis. Image Underst.3
2018 Incremental and Iterative Learning of Answer Set Programs from Mutually Distinct Examples
abstract
Abstract Over the years the Artificial Intelligence (AI) community has produced several datasets which have given the machine learning algorithms the opportunity to learn various skills across various domains. However, a subclass of these machine learning algorithms that aimed at learning logic programs, namely the Inductive Logic Programming algorithms, have often failed at the task due to the vastness of these datasets. This has impacted the usability of knowledge representation and reasoning techniques in the development of AI systems. In this research, we try to address this scalability issue for the algorithms that learn answer set programs. We present a sound and complete algorithm which takes the input in a slightly different manner and performs an efficient and more user controlled search for a solution. We show via experiments that our algorithm can learn from two popular datasets from machine learning community, namely bAbl (a question answering dataset) and MNIST (a dataset for handwritten digit recognition), which to the best of our knowledge was not previously possible. The system is publicly available at https://goo.gl/KdWAcV .
Arindam Mitra, Chitta Baral
Theory Pract. Log. Program.2
2017 Revision and Updates in Possibly Action-Occurrence-Incomplete Narratives
Chitta Baral, Tran Cao Son
PRIMA1
2016 Addressing a Question Answering Challenge by Combining Statistical Methods with Inductive Rule Learning and Reasoning
abstract
A group of researchers from Facebook has recently proposed a set of 20 question-answering tasks (Facebook's bAbl dataset) as a challenge for the natural language understanding ability of an intelligent agent. These tasks are designed to measure various skills of an agent, such as: fact based question-answering, simple induction, the ability to find paths, co-reference resolution and many more. Their goal is to aid in the development of systems that can learn to solve such tasks and to allow a proper evaluation of such systems. They show existing systems cannot fully solve many of those toy tasks. In this work, we present a system that excels at all the tasks except one. The proposed model of the agent uses the Answer Set Programming (ASP) language as the primary knowledge representation and reasoning language along with the standard statistical Natural Language Processing (NLP) models. Given a training dataset containing a set of narrations, questions and their answers, the agent jointly uses a translation system, an Inductive Logic Programming algorithm and Statistical NLP methods to learn the knowledge needed to answer similar questions. Our results demonstrate that the introduction of a reasoning module significantly improves the performance of an intelligent agent.
Arindam Mitra, Chitta Baral
AAAI2
2016 Learning To Use Formulas To Solve Simple Arithmetic Problems
abstract
Solving simple arithmetic word problems is one of the challenges in Natural Language Understanding.This paper presents a novel method to learn to use formulas to solve simple arithmetic word problems.Our system, analyzes each of the sentences to identify the variables and their attributes; and automatically maps this information into a higher level representation.It then uses that representation to recognize the presence of a formula along with its associated variables.An equation is then generated from the formal description of the formula.In the training phase, it learns to score the pair from the systematically generated higher level representation.It is able to solve 86.07% of the problems in a corpus of standard primary school test questions and beats the state-of-the-art by a margin of 8.07%.
Arindam Mitra, Chitta Baral
ACL (1)2
2016 Plan Failure Analysis: Formalization and Application in Interactive Planning Through Natural Language Communication
Chitta Baral, Tran Cao Son, Michael Gelfond, Arindam Mitra
PRIMA1
2015 Knowledge Representation and Reasoning: What's Hot
abstract
This is an extended abstract about what is hot in the field of Knowledge Representation and Reasoning.
Chitta Baral, Giuseppe De Giacomo
AAAI1
2015 Exploring the KD45 Property of a Kripke Model After the Execution of an Action Sequence
abstract
The paper proposes a condition for preserving the KD45 property of a Kripke model when a sequence of update models is applied to it. The paper defines the notions of a primitive update model and a semi-reflexive KD45 (or sr-KD45) Kripke model. It proves that updating a sr-KD45 Kripke model using a primitive update model results in a sr-KD45 Kripke model, i.e., a primitive update model preserves the properties of a sr-KD45 Kripke model. It shows that several update models for modeling well-known actions found in the literature are primitive. This result provides guarantees that can be useful in presence of multiple applications of actions in multi-agent system (e.g., multi-agent planning).
Tran Cao Son, Enrico Pontelli, Chitta Baral, Gregory Gelfond
AAAI3
2015 The NL2KR Platform for building Natural Language Translation Systems
abstract
Nguyen Vo, Arindam Mitra, Chitta Baral. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Nguyen Ha Vo, Arindam Mitra, Chitta Baral
ACL (1)3
2015 Learning to Automatically Solve Logic Grid Puzzles
abstract
Logic grid puzzle is a genre of logic puzzles in which we are given (in a natural language) a scenario, the object to be deduced and certain clues.The reader has to figure out the solution using the clues provided and some generic domain constraints.In this paper, we present a system, LOGICIA, that takes a logic grid puzzle and the set of elements in the puzzle and tries to solve it by translating it to the knowledge representation and reasoning language of Answer Set Programming (ASP) and then using an ASP solver.The translation to ASP involves extraction of entities and their relations from the clues.For that we use a novel learning based approach which uses varied supervision, including the entities present in a clue and the expected representation of a clue in ASP.Our system, LO-GICIA, learns to automatically translate a clue with 81.11% accuracy and is able to solve 71% of the problems of a corpus.This is the first learning system that can solve logic grid puzzles described in natural language in a fully automated manner.The code and the data will be made publicly available at http://bioai.lab.asu.edu/logicgridpuzzles.
Arindam Mitra, Chitta Baral
EMNLP2
2015 Towards Addressing the Winograd Schema Challenge - Building and Using a Semantic Parser and a Knowledge Hunting Module
Arpit Sharma 0001, Nguyen Ha Vo, Somak Aditya, Chitta Baral
IJCAI4
2015 "Add Another Blue Stack of the Same Height!": ASP Based Planning and Plan Failure Analysis
Chitta Baral, Tran Cao Son
LPNMR1
2015 Recognizing Social Constructs from Textual Conversation
abstract
Somak Aditya, Chitta Baral, Nguyen Ha Vo, Joohyung Lee, Jieping Ye, Zaw Naung, Barry Lumpkin, Jenny Hastings, Richard Scherl, Dawn M. Sweet, Daniela Inclezan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Somak Aditya, Chitta Baral, Nguyen Ha Vo, Jieping Ye, Zaw Naung, Barry Lumpkin, Jenny Hastings, Richard B. Scherl, Dawn M. Sweet, Daniela Inclezan
HLT-NAACL2
2014 Pathway Specification and Comparative Queries: A High Level Language with Petri Net Semantics
abstract
Understanding biological pathways is an important activity in the biological domain for drug development. Due to the parallelism and complexity inherent in pathways, computer models that can answer queries about pathways are needed. A researcher may ask `what-if' questions comparing alternate scenarios, that require deeper understanding of the underlying model. In this paper, we present overview of such a system we developed and an English-like high level language to express pathways and queries. Our language is inspired by high level action and query languages and it uses Petri Net execution semantics.
Saadat Anwar, Chitta Baral
AAAI2
2014 Finitary S5-Theories
Tran Cao Son, Enrico Pontelli, Chitta Baral, Gregory Gelfond
JELIA3
2013 Encoding Higher Level Extensions of Petri Nets in Answer Set Programming
Saadat Anwar, Chitta Baral, Katsumi Inoue
LPNMR2
2013 Event-Object Reasoning with Curated Knowledge Bases: Deriving Missing Information
Chitta Baral, Nguyen Ha Vo
LPNMR1
2013 Encoding Petri Nets in Answer Set Programming for Simulation Based Reasoning
Saadat Anwar, Chitta Baral, Katsumi Inoue
Theory Pract. Log. Program.2
2012 Solving Puzzles Described in English by Automated Translation to Answer Set Programming and Learning How to Do that Translation
Chitta Baral, Juraj Dzifcak
KR1
2012 From Knowledge Represented in Frame-Based Languages to Declarative Representation and Reasoning via ASP
Chitta Baral, Shanshan Liang
KR1
2012 A SNPshot of PubMed to associate genetic variants with drugs, diseases, and adverse reactions
Jörg Hakenberg, Dmitry Voronov, Nguyen Ha Vo, Shanshan Liang, Saadat Anwar, Barry Lumpkin, Robert Leaman, Luis Tari, Chitta Baral
J. Biomed. Informatics9
2012 Incremental Information Extraction Using Relational Databases
abstract
Information extraction systems are traditionally implemented as a pipeline of special-purpose processing modules targeting the extraction of a particular kind of information. A major drawback of such an approach is that whenever a new extraction goal emerges or a module is improved, extraction has to be reapplied from scratch to the entire text corpus even though only a small part of the corpus might be affected. In this paper, we describe a novel approach for information extraction in which extraction needs are expressed in the form of database queries, which are evaluated and optimized by database systems. Using database queries for information extraction enables generic extraction and minimizes reprocessing of data by performing incremental extraction to identify which part of the data is affected by the change of components or goals. Furthermore, our approach provides automated query generation components so that casual users do not have to learn the query language in order to perform extraction. To demonstrate the feasibility of our incremental extraction approach, we performed experiments to highlight two important aspects of an information extraction system: efficiency and quality of extraction results. Our experiments show that in the event of deployment of a new module, our incremental extraction approach reduces the processing time by 89.64 percent as compared to a traditional pipeline approach. By applying our methods to a corpus of 17 million biomedical abstracts, our experiments show that the query performance is efficient for real-time applications. Our experiments also revealed that our approach achieves high quality extraction results.
Luis Tari, Phan Huy Tu, Jörg Hakenberg, Yi Chen 0001, Tran Cao Son, Graciela Gonzalez-Hernandez, Chitta Baral
IEEE Trans. Knowl. Data Eng.7
2012 Typed answer set programming lambda calculus theories and correctness of inverse lambda algorithms with respect to them
abstract
Abstract Our broader goal is to automatically translate English sentences into formulas in appropriate knowledge representation languages as a step towards understanding and thus answering questions with respect to English text. Our focus in this paper is on the language of Answer Set Programming (ASP). Our approach to translate sentences to ASP rules is inspired by Montague's use of lambda calculus formulas as meaning of words and phrases. With ASP as the target language the meaning of words and phrases are ASP-lambda formulas. In an earlier work we illustrated our approach by manually developing a dictionary of words and their ASP-lambda formulas. However such an approach is not scalable. In this paper our focus is on two algorithms that allow one to construct ASP-lambda formulas in an inverse manner. In particular the two algorithms take as input two lambda-calculus expressions G and H and compute a lambda-calculus expression F such that F with input as G, denoted by F@G, is equal to H; and similarly G@F = H. We present correctness and complexity results about these algorithms. To do that we develop the notion of typed ASP-lambda calculus theories and their orders and use it in developing the completeness results.
Chitta Baral, Juraj Dzifcak, Marcos Alvarez Gonzalez, Aaron Gottesman
Theory Pract. Log. Program.1
2011 Lessons from Efforts to Automatically Translate English to Knowledge Representation Languages
Chitta Baral
LPNMR1
2011 Molecular Event Extraction from Link Grammar Parse Trees in the BioNLP'09 Shared Task
abstract
The BioNLP’09 Shared Task deals with extracting information on molecular events, such as gene expression and protein localization, from natural language text. Information in this benchmark are given as tuples including protein names, trigger terms for each event, and possible other participants such as bindings sites. We address all three tasks of BioNLP’09: event detection, event enrichment, and recognition of negation and speculation. Our method for the first two tasks is based on a deep parser; we store the parse tree of each sentence in a relational database scheme. From the training data, we collect the dependencies connecting any two relevant terms of a known tuple, that is, the shortest paths linking these two constituents. We encode all such linkages in a query language to retrieve similar linkages from unseen text. For the third task, we rely on a hierarchy of hand-crafted regular expressions to recognize speculation and negated events. In this paper, we added extensions regarding a post-processing step that handles ambiguous event trigger terms, as well as an extension of the query language to relax linkage constraints. On the BioNLP Shared Task test data, we achieve an overall F1-measure of 32%, 29%, and 30% for the successive Tasks 1, 2, and 3, respectively.
Jörg Hakenberg, Illés Solt, Domonkos Tikk, Nguyen Ha Vo, Luis Tari, Quang Long Nguyen, Chitta Baral, Ulf Leser
Comput. Intell.7
2010 GenerIE: Information extraction using database queries
abstract
Information extraction systems are traditionally implemented as a pipeline of special-purpose processing modules. A major drawback of such an approach is that whenever a new extraction goal emerges or a module is improved, extraction has to be re-applied from scratch to the entire text corpus even though only a small part of the corpus might be affected. In this demonstration proposal, we describe a novel paradigm for information extraction: we store the parse trees output by text processing in a database, and then express extraction needs using queries, which can be evaluated and optimized by databases. Compared with the existing approaches, database queries for information extraction enable generic extraction and minimize reprocessing. However, such an approach also poses a lot of technical challenges, such as language design, optimization and automatic query generation. We will present the opportunities and challenges that we met when building GenerIE, a system that implements this paradigm.
Luis Tari, Phan Huy Tu, Jörg Hakenberg, Yi Chen 0001, Tran Cao Son, Graciela Gonzalez-Hernandez, Chitta Baral
ICDE7
2010 Reasoning about Actions and Change: From Single Agent Actions to Multi-Agent Actions (Extended Abstract)
Chitta Baral
KR1
2010 Invited Presentations at the Twelfth International Conference on Principles of Knowledge Representation and Reasoning
Chitta Baral, Ian Horrocks 0001, Yoav Shoham
KR1
2010 Discovering drug-drug interactions: a text-mining and reasoning approach based on properties of drug metabolism
abstract
MOTIVATION: Identifying drug-drug interactions (DDIs) is a critical process in drug administration and drug development. Clinical support tools often provide comprehensive lists of DDIs, but they usually lack the supporting scientific evidences and different tools can return inconsistent results. In this article, we propose a novel approach that integrates text mining and automated reasoning to derive DDIs. Through the extraction of various facts of drug metabolism, not only the DDIs that are explicitly mentioned in text can be extracted but also the potential interactions that can be inferred by reasoning. RESULTS: Our approach was able to find several potential DDIs that are not present in DrugBank. We manually evaluated these interactions based on their supporting evidences, and our analysis revealed that 81.3% of these interactions are determined to be correct. This suggests that our approach can uncover potential DDIs with scientific evidences explaining the mechanism of the interactions.
Luis Tari, Saadat Anwar, Shanshan Liang, James Cai, Chitta Baral
Bioinform.5
2010 Efficient Extraction of Protein-Protein Interactions from Full-Text Articles
abstract
Proteins and their interactions govern virtually all cellular processes, such as regulation, signaling, metabolism, and structure. Most experimental findings pertaining to such interactions are discussed in research papers, which, in turn, get curated by protein interaction databases. Authors, editors, and publishers benefit from efforts to alleviate the tasks of searching for relevant papers, evidence for physical interactions, and proper identifiers for each protein involved. The BioCreative II.5 community challenge addressed these tasks in a competition-style assessment to evaluate and compare different methodologies, to make aware of the increasing accuracy of automated methods, and to guide future implementations. In this paper, we present our approaches for protein-named entity recognition, including normalization, and for extraction of protein-protein interactions from full text. Our overall goal is to identify efficient individual components, and we compare various compositions to handle a single full-text article in between 10 seconds and 2 minutes. We propose strategies to transfer document-level annotations to the sentence-level, which allows for the creation of a more fine-grained training corpus; we use this corpus to automatically derive around 5,000 patterns. We rank sentences by relevance to the task of finding novel interactions with physical evidence, using a sentence classifier built from this training corpus. Heuristics for paraphrasing sentences help to further remove unnecessary information that might interfere with patterns, such as additional adjectives, clauses, or bracketed expressions. In BioCreative II.5, we achieved an f-score of 22 percent for finding protein interactions, and 43 percent for mapping proteins to UniProt IDs; disregarding species, f-scores are 30 percent and 55 percent, respectively. On average, our best-performing setup required around 2 minutes per full text. All data and pattern sets as well as Java classes that extend- - third-party software are available as supplementary information (see Appendix).
Jörg Hakenberg, Robert Leaman, Nguyen Ha Vo, Siddhartha Jonnalagadda, Ryan Sullivan, Luis Tari, Chitta Baral, Graciela Gonzalez-Hernandez
IEEE ACM Trans. Comput. Biol. Bioinform.8
2010 Logic programming for finding models in the logics of knowledge and its applications: A case study
abstract
Abstract The logics of knowledge are modal logics that have been shown to be effective in representing and reasoning about knowledge in multi-agent domains. Relatively few computational frameworks for dealing with computation of models and useful transformations in logics of knowledge (e.g., to support multi-agent planning with knowledge actions and degrees of visibility) have been proposed. This paper explores the use of logic programming (LP) to encode interesting forms of logics of knowledge and compute Kripke models. The LP modeling is expanded with useful operators on Kripke structures, to support multi-agent planning in the presence of both world-altering and knowledge actions. This results in the first ever implementation of a planner for this type of complex multi-agent domains.
Chitta Baral, Gregory Gelfond, Enrico Pontelli, Tran Cao Son
Theory Pract. Log. Program.1
2009 What to do and how to do it: Translating natural language directives into temporal and dynamic logic representation for goal management and action execution
abstract
Robots that can be given instructions in spoken language need to be able to parse a natural language utterance quickly, determine its meaning, generate a goal representation from it, check whether the new goal conflicts with existing goals, and if acceptable, produce an action sequence to achieve the new goal (ideally being sensitive to the existing goals). In this paper, we describe an integrated robotic architecture that can achieve the above steps by translating natural language instructions incrementally and simultaneously into formal logical goal description and action languages, which can be used both to reason about the achievability of a goal as well as to generate new action scripts to pursue the goal. We demonstrate the implementation of our approach on a robot taking spoken natural language instructions in an office environment.
Juraj Dzifcak, Matthias Scheutz, Chitta Baral, Paul W. Schermerhorn
ICRA3
2009 Modeling Multi-agent Domains in an Action Languages: An Empirical Study Using
Chitta Baral, Tran Cao Son, Enrico Pontelli
LPNMR1
2009 Fuzzy c-means clustering with prior biological knowledge
Luis Tari, Chitta Baral, Seungchan Kim
J. Biomed. Informatics2
2009 Probabilistic reasoning with answer sets
abstract
Abstract This paper develops a declarative language, P-log, that combines logical and probabilistic arguments in its reasoning. Answer Set Prolog is used as the logical foundation, while causal Bayes nets serve as a probabilistic foundation. We give several non-trivial examples and illustrate the use of P-log for knowledge representation and updating of knowledge. We argue that our approach to updates is more appealing than existing approaches. We give sufficiency conditions for the coherency of P-log programs and show that Bayes nets can be easily mapped to coherent P-log programs.
Chitta Baral, Michael Gelfond, J. Nelson Rushton
Theory Pract. Log. Program.1
2008 Using Answer Set Programming and Lambda Calculus to Characterize Natural Language Sentences with Normatives and Exceptions
Chitta Baral, Juraj Dzifcak, Tran Cao Son
AAAI1
2008 Non-monotonic Temporal Logics that Facilitate Elaboration Tolerant Revision of Goals
Chitta Baral, Jicheng Zhao
AAAI1
2008 Extracting Protein-Protein Interactions from MEDLINE Using Syntactic Roles
abstract
With rapid growth in genomics research in last decade, amount of information a biomedical researcher has to keep track of and understand has increased tremendously. We present a fully automated information extraction system to aid these researchers to identify and locate gene and protein interactions in biomedical text. Our extraction system handles complex sentences and extracts multiple and nested interactions specified in these sentences. Experimental evaluations with two other state of the art extraction systems indicate that the IntEx system achieves better performance without the labor intensive pattern engineering requirement.
Syed Toufeeq Ahmed, Hasan Davulcu, Chitta Baral
BIBM3
2008 Using Answer Set Programming for Knowledge Representation and Reasoning: Future Directions
Chitta Baral
ICLP1
2008 State-Based Regression with Sensing and Knowledge
Richard B. Scherl, Tran Cao Son, Chitta Baral
PRICAI3
2008 Maintenance goals of agents in a dynamic environment: Formulation and policy construction
Chitta Baral, Thomas Eiter, Marcus Bjäreland, Mutsumi Nakamura
Artif. Intell.1
2007 Towards Overcoming the Knowledge Acquisition Bottleneck in Answer Set Prolog Applications: Embracing Natural Language Inputs
Chitta Baral, Juraj Dzifcak, Luis Tari
ICLP1
2007 Using the Probabilistic Logic Programming Language P-log for Causal and Counterfactual Reasoning and Non-Naive Conditioning
Chitta Baral, Matt Hunsaker
IJCAI1
2007 Non-monotonic Temporal Logics for Goal Specification
Chitta Baral, Jicheng Zhao
IJCAI1
2007 Reasoning and planning with sensing actions, incomplete information, and static causal laws using answer set programming
abstract
Abstract We extend the 0-approximation of sensing actions and incomplete information in Son and Baral (2001) to action theories with static causal laws and prove its soundness with respect to the possible world semantics. We also show that the conditional planning problem with respect to this approximation isNP-complete. We then present an answer set programming based conditional planner, called ASCP, that is capable of generating both conformant plans and conditional plans in the presence of sensing actions, incomplete information about the initial state, and static causal laws. We prove the correctness of our implementation and argue that our planner is sound and complete with respect to the proposed approximation. Finally, we present experimental results comparing ASCP to other planners.
Phan Huy Tu, Tran Cao Son, Chitta Baral
Theory Pract. Log. Program.3
2006 Goal Specification, Non-Determinism and Quantifying over Policies
Chitta Baral, Jicheng Zhao
AAAI1
2006 Macros, Macro Calls and Use of Ensembles in Modular Answer Set Programming
Chitta Baral, Juraj Dzifcak
ICLP1
2006 A State-Based Regression Formulation for Domains with Sensing Actions and Incomplete Information
abstract
We present a state-based regression function for planning domains where an agent does not have complete information and may have sensing actions. We consider binary domains and employ a three-valued characterization of domains with sensing actions to define the regression function. We prove the soundness and completeness of our regression formulation with respect to the definition of progression. More specifically, we show that (i) a plan obtained through regression for a planning problem is indeed a progression solution of that planning problem, and that (ii) for each plan found through progression, using regression one obtains that plan or an equivalent one.
Le-Chi Tuan, Chitta Baral, Tran Cao Son
Log. Methods Comput. Sci.2
2006 Domain-dependent knowledge in answer set planning
abstract
In this article we consider three different kinds of domain-dependent control knowledge (temporal, procedural and HTN-based) that are useful in planning. Our approach is declarative and relies on the language of logic programming with answer set semantics (AnsProlog*). AnsProlog* is designed to plan without control knowledge. We show how temporal, procedural and HTN-based control knowledge can be incorporated into AnsProlog* by the modular addition of a small number of domain-dependent rules, without the need to modify the planner. We formally prove the correctness of our planner, both in the absence and presence of the control knowledge. Finally, we perform some initial experimentation that demonstrates the potential reduction in planning time that can be achieved when procedural domain knowledge is used to solve planning problems with large plan length.
Tran Cao Son, Chitta Baral, Tran Hoai Nam, Sheila A. McIlraith
ACM Trans. Comput. Log.2
2005 Using SAT and Logic Programming to Design Polynomial-Time Algorithms for Planning in Non-Deterministic Domains
Chitta Baral, Thomas Eiter, Jicheng Zhao
AAAI1
2005 Reasoning about Intended Actions
Chitta Baral, Michael Gelfond
AAAI1
2005 Issues in Reasoning about Interaction Networks in Cells: Necessity of Event Ordering Knowledge
Tran Hoai Nam, Chitta Baral, Carron Shankland
AAAI2
2005 An Algorithm to Learn Causal Relations Between Genes from Steady State Data: Simulation and Its Application to Melanoma Dataset
Xin Zhang 0005, Chitta Baral, Seungchan Kim
AIME2
2005 Knowledge updates: Semantics and complexity issues
Chitta Baral, Yan Zhang 0003
Artif. Intell.1
2004 Adding Time and Intervals to Procedural and Hierarchical Control Specifications
Tran Cao Son, Chitta Baral, Le-Chi Tuan
AAAI2
2004 Encoding Probabilistic Causal Model in Probabilistic Action Language
Tran Hoai Nam, Chitta Baral
AAAI2
2004 Regression with Respect to Sensing Actions and Partial States
Le-Chi Tuan, Chitta Baral, Xin Zhang 0005, Tran Cao Son
AAAI2
2004 Goal Specification in Presence of Non-Deterministic Actions
Chitta Baral, Jicheng Zhao
ECAI1
2004 A Polynomial-Time Algorithm for Constructing k-Maintainable Policies
Chitta Baral, Thomas Eiter
KR1
2004 Reasoning about Triggered Actions in AnsProlog and Its Application to Molecular Interactions in Cells
Tran Hoai Nam, Chitta Baral
KR2
2004 Probabilistic Reasoning With Answer Sets
Chitta Baral, Michael Gelfond, J. Nelson Rushton
LPNMR1
2004 Planning with Sensing Actions and Incomplete Information Using Logic Programming
Tran Cao Son, Phan Huy Tu, Chitta Baral
LPNMR3
2003 Introduction to the special issue on Programming with Answer Sets
abstract
The search for an appropriate characterization of negation as failure in logic programs in the mid 1980s led to several proposals. Amongst them the stable model semantics – later referred to as answer set semantics, and the well-founded semantics are the most popular and widely referred ones. According to the latest (September 2002) list of most cited source documents in the CiteSeer database (http://citeseer.nj.nec.com) the original stable model semantics paper (Gelfond and Lifschitz, 1988) is ranked 10th with 649 citations and the well-founded semantics paper (Van Gelder et al., 1991) is ranked 70th with 306 citations. Since 1988 – when stable models semantics was proposed – there has been a large body of work centered around logic programs with answer set semantics covering topics such as: systematic program development, systematic program analysis, knowledge representation, declarative problem solving, answer set computing algorithms, complexity and expressiveness, answer set computing systems, relation with other non-monotonic and knowledge representation formalisms, and applications to various tasks.
Chitta Baral, Alessandro Provetti, Tran Cao Son
Theory Pract. Log. Program.1
2002 A Transition Function Based Characterization of Actions with Delayed and Continuous Effects
Chitta Baral, Tran Cao Son, Le-Chi Tuan
KR1
2002 The Complexity of Model Checking for Knowledge Update
Chitta Baral
KR1
2001 Computational Complexity of Planning with Temporal Goals
Chitta Baral, Vladik Kreinovich, Raul Trejo
IJCAI1
2001 On the Semantics of Knowledge Update
Chitta Baral
IJCAI1
2001 Declarative Specification and Solution of Combinatorial Auctions Using Logic Programming
Chitta Baral, Cenk Uyan
LPNMR1
2001 Planning with Different Forms of Domain-Dependent Control Knowledge - An Answer Set Programming Approach
Tran Cao Son, Chitta Baral, Sheila A. McIlraith
LPNMR2
2001 Formalizing sensing actions A transition function based approach
Tran Cao Son, Chitta Baral
Artif. Intell.2
2001 Formalizing and Reasoning About the Requirements Specifications of Workflow Systems
abstract
This work addresses the problem of workflow requirements specifications considering the realistic assumptions that, it involves experts from different domains (i.e. representatives of different business policies); not all the possible execution scenarios are known beforehand, during the early stage of specification. In particular, since the main purpose of a workflow is to achieve a certain (bussiness) goal, we propose a formalism which enables the users to specify their requirements (and expectations) and test if the information that they have provided is, in a sense, sufficient for the workflow to behave "as desired", in terms of the goal. Our methodology allows domain experts to express not only their knowledge, but also the "ignorance" (the semantics allows for unknown values to reflect a realistic situation of agents dealing with incomplete information) and the possibility of occurrence of exceptional situations. As a basis for formalizing the process of equirements specifications, we are using the recent results on reasoning about actions. We propose a high level language AW which enables specifying the effects that activites have on the environment and how they should be coordinated. We also describe our prototype tool for process specification. Strictly speaking, in this work we go "one step" before actual analysis and design, and offer a formalism which enables the involved partners to see if the extent to which they have expressed their domain knowledge (which may sometimes be subject to a proprietary restricions) can satisfy the intended needs and behaviour of their product_to_be. We define an entailment relation which enables reasoning about the correctness of the specification, in terms of achieving a desired goal and, also testing about consequences of modifications in the workflow descriptions.
Goce Trajcevski, Chitta Baral, Jorge Lobo 0001
Int. J. Cooperative Inf. Syst.2
2001 From Planning to Searching for the Shortest Plan: An Optimal Transition
Raul Trejo, Joel Galloway, Charanjiv Sachar, Vladik Kreinovich, Chitta Baral, Le-Chi Tuan
Int. J. Uncertain. Fuzziness Knowl. Based Syst.5
2000 Formalizing (and Reasoning About) the Specifications of Workflows
Goce Trajcevski, Chitta Baral, Jorge Lobo 0001
CoopIS2
2000 Formulating diagnostic problem solving using an action language with narratives and sensing
Chitta Baral, Sheila A. McIlraith, Tran Cao Son
KR1
2000 Abductive reasoning through filtering
Chitta Baral
Artif. Intell.1
2000 Computational complexity of planning and approximate planning in the presence of incompleteness
Chitta Baral, Vladik Kreinovich, Raul Trejo
Artif. Intell.1
1999 Computational Complexity of Planning and Approximate Planning in Presence of Incompleteness
Chitta Baral, Vladik Kreinovich, Raul Trejo
IJCAI1
1998 Design and Implementation of Display Specification for Multimedia Answers
abstract
We present the design and implementation of a loosely-bound SQL extension that allows users to include high-level display specifications with an SQL query, particularly when dealing with multimedia databases. We describe an architecture that allows a relatively simple implementation of dynamic query browsers using the proposed query language on stand-alone applications or World Wide Web pages. We have already implemented most of our proposed extension.
Chitta Baral, Graciela Gonzalez-Hernandez, Tran Cao Son
ICDE1
1998 SQL+D: Extended Display Capabilities for Multimedia Database Queries
Chitta Baral, Graciela Gonzalez-Hernandez, Amarendra Nandigam
ACM Multimedia1
1998 Value Minimization in Circumscription
Chitta Baral, Alfredo Gabaldon, Alessandro Provetti
Artif. Intell.1
1998 Formalizing Narratives Using Nested Circumscription
Chitta Baral, Alfredo Gabaldon, Alessandro Provetti
Artif. Intell.1
1998 Conceptual Modeling and Querying in Multimedia Databases
Chitta Baral, Graciela Gonzalez-Hernandez, Tran Cao Son
Multim. Tools Appl.1
1997 Defeasible Specifications in Action Theories
Chitta Baral, Jorge Lobo 0001
IJCAI1
1996 Value Minimization in Circumscription
Chitta Baral, Alfredo Gabaldon, Alessandro Provetti
KR1
1995 Reasoning about actions: Non-deterministic effects, Constraints, and Qualification
Chitta Baral
IJCAI1
1994 Rule Based Updates on Simple Knowledge Bases
Chitta Baral
AAAI1
1994 Varying Selection Functions to Relate Conditional Logics and Preferential Models
abstract
In this paper we study the relationship between preferential model semantics of Shoham [19] and Kraus et al. [12] with the flat version of conditional logic denned by Bell [3] and its variations. We present a new selection function without constraini
Chitta Baral
Fundam. Informaticae1
1994 Combining Default Logic Databases
abstract
During the past decade, it has become increasingly clear that the future generation of large-scale knowledge bases will consist, not of one single isolated knowledge base, but a multiplicity of specialized knowledge bases that contain knowledge about different domains of expertise. These knowledge bases will work cooperatively, pooling together their varied bodies of knowledge, so as to be able to solve complex problems that no single knowledge base, by itself, would have been able to address successfully. In any such situation, inconsistencies are bound to arise. In this paper, we address the question: "Suppose we have a set of knowledge bases, KB1, …, KBn, each of which uses default logic as the formalism for knowledge representation, and a set of integrity constraints IC. What knowledge base constitutes an acceptable combination of KB1, …, KBn?"
Chitta Baral, Sarit Kraus, Jack Minker, V. S. Subrahmanian
Int. J. Cooperative Inf. Syst.1
1993 Representing Concurrent Actions in Extended Logic Programming
Chitta Baral, Michael Gelfond
IJCAI1
1993 Dualities Between Alternative Semantics for Logic Programming and Nonmonotonic Reasoning
Chitta Baral, V. S. Subrahmanian
J. Autom. Reason.1
1992 Generalized Negation As Failure and Semantics of Normal Disjunctive Logic Programs
Chitta Baral
LPAR1
1992 Combining Knowledge Bases Consisting of First-Order Analysis
abstract
Consider the construction of an expert system by encoding the knowledge of different experts. Suppose the knowledge provided by each expert is encoded into a knowledge base. Then the process ofcombiningthe knowledge of these different experts is an important and nontrivial problem. We study this problem here when the expert systems are considered to be first‐order theories. We present techniques for resolving inconsistencies in such knowledge bases. We also provide algorithms for implementing these techniques.
Chitta Baral, Sarit Kraus, Jack Minker, V. S. Subrahmanian
Comput. Intell.1
1992 Stable and Extension Class Theory for Logic Programs and Default Logics
Chitta Baral, V. S. Subrahmanian
J. Autom. Reason.1
1991 Combining Knowledge Bases Consisting of First Order Theories
Chitta Baral, Sarit Kraus, Jack Minker, V. S. Subrahmanian
ISMIS1
1991 WF³: A Semantics for Negation in Normal Disjunctive Logic Programs
Chitta Baral, Jorge Lobo 0001, Jack Minker
ISMIS1
1991 Combining Multiple Knowledge Bases
abstract
Combining knowledge present in multiple knowledge base systems into a single knowledge base is discussed. A knowledge based system can be considered an extension of a deductive database in that it permits function symbols as part of the theory. Alternative knowledge bases that deal with the same subject matter are considered. The authors define the concept of combining knowledge present in a set of knowledge bases and present algorithms to maximally combine them so that the combination is consistent with respect to the integrity constraints associated with the knowledge bases. For this, the authors define the concept of maximality and prove that the algorithms presented combine the knowledge bases to generate a maximal theory. The authors also discuss the relationships between combining multiple knowledge bases and the view update problem.>
Chitta Baral, Sarit Kraus, Jack Minker
IEEE Trans. Knowl. Data Eng.1
1990 Generalized Well-founded Semantics for Logic Programs (Extended Abstract)
Chitta Baral, Jorge Lobo 0001, Jack Minker
CADE1