Vivek Gupta 0001

dblp:71/5332-1 · DBLP profile ↗
← Back
27ranked-venue papers
9as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 REaR : Retrieve, Expand and Refine for Effective Multitable Retrieval
abstract
Rishita Agarwal, Himanshu Singhal, Peter Baile Chen, Manan Roy Choudhury, Dan Roth, Vivek Gupta. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Rishita Agarwal, Himanshu Singhal, Peter Baile Chen, Manan Roy Choudhury, Dan Roth 0001, Vivek Gupta 0001
ACL (1)6
2026 TabReX: Tabular Referenceless eXplainable Evaluation
abstract
Evaluating the quality of tables generated by large language models (LLMs) remains an open challenge: existing metrics either flatten tables into text, ignoring structure, or rely on fixed references that limit generalization.We present TABREX, a reference-less, propertydriven framework for evaluating tabular generation via graph-based reasoning.TABREX converts both source text and generated tables into canonical knowledge graphs, aligns them through an LLM-guided matching process, and computes interpretable, rubric-aware scores that quantify structural and factual fidelity.The resulting metric provides controllable tradeoffs between sensitivity and specificity, yielding human-aligned judgments and cell-level error traces.To systematically assess metric robustness, we introduce TABREX-BENCH, a large-scale benchmark spanning six domains and twelve planner-driven perturbation types across three difficulty tiers.Empirical results show that TABREX achieves the highest correlation with expert rankings, remains stable under harder perturbations, and enables finegrained model-vs-prompt analysis establishing a new paradigm for trustworthy, explainable evaluation of structured generation systems.
Tejas Anvekar, Junha Park, Aparna Garimella, Vivek Gupta 0001
ACL (1)4
2026 The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
abstract
Recent advances in multimodal large language models (MLLMs) have yielded increasingly powerful models, yet their perceptual capacities remain poorly characterized. In practice, most model families scale language component while reusing nearly identical vision encoders (e.g., Qwen2.5-VL 3B/7B/72B), which raises pivotal concerns about whether progress reflects genuine visual grounding or reliance on internet-scale textual world knowledge. Existing evaluation methods emphasize end-task accuracy, overlooking robustness, attribution fidelity, and reasoning under controlled perturbations. We present The Perceptual Observatory, a framework that characterizes MLLMs across verticals like: (i) simple vision tasks, such as face matching and text-in-vision comprehension capabilities; (ii) local-to-global understanding, encompassing image matching, grid pointing game, and attribute localization, which tests general visual grounding. Each vertical is instantiated with ground-truth datasets of faces and words, systematically perturbed through pixel-based augmentations and diffusion-based stylized illusions. The Perceptual Observatory moves beyond leaderboard accuracy to yield insights into how MLLMs preserve perceptual grounding and relational structure under perturbations, providing a principled foundation for analyzing strengths and weaknesses of current and future models.
Tejas Anvekar, Fenil Denish Bardoliya, Pavan Turaga, Chitta Baral, Vivek Gupta 0001
WACV5
2026 MapVerse: A Benchmark for Geospatial Question Answering on Diverse Real-World Maps
abstract
Maps are powerful carriers of structured and contextual knowledge, encompassing geography, demographics, infrastructure, and environmental patterns. Reasoning over such knowledge requires models to integrate spatial relationships, visual cues, real-world context, and domain-specific expertise—capabilities that current large language models (LLMs) and vision–language models (VLMs) still struggle to exhibit consistently. Yet, datasets used to benchmark VLMs on map-based reasoning remain narrow in scope, restricted to specific domains, and heavily reliant on artificially generated content (outputs from LLMs or pipeline-based methods), offering limited depth for evaluating genuine geospatial reasoning. To address this gap, we present MAPVERSE, a large-scale benchmark built on real-world maps. It comprises 11,837 human-authored question-answer pairs across 1,025 maps, spanning ten diverse map categories and multiple question categories for each. The dataset provides a rich setting for evaluating map reading, interpretation, and multimodal reasoning. We evaluate ten state-of-the-art models against our benchmark to establish baselines and quantify reasoning gaps. Beyond overall performance, we conduct fine-grained categorical analyses to assess model inference across multiple dimensions and investigate the visual factors shaping reasoning outcomes. Our findings reveal that while current VLMs perform competitively on classification-style tasks, both open-and closed-source models fall short on advanced tasks requiring complex spatial reasoning.
Sharat Bhat, Harshita Khandelwal, Tushar Kataria, Vivek Gupta 0001
WACV4
2025 Map&Make: Schema Guided Text to Table Generation
abstract
Transforming dense, unstructured text into interpretable tables-commonly referred to as Text-to-Table generation-is a key task in information extraction.Existing methods often overlook what complex information to extract and how to infer it from text.We present Map&Make, a versatile approach that decomposes text into atomic propositions to infer latent schemas, which are then used to generate tables capturing both qualitative nuances and quantitative facts.We evaluate our method on three challenging datasets: Rotowire, known for its complex, multi-table schema; Livesum which requires numerical aggregation; and Wiki40 which require open text extraction from mulitple domains.By correcting hallucination errors in Rotowire, we also provide a cleaner benchmark.Our method shows significant gains in both accuracy and interpretability across comprehensive comparative and referenceless metrics.Finally, ablation studies highlight the key factors driving performance and validate the utility of our approach in structured summarization.Code and data are available
Naman Ahuja, Fenil Denish Bardoliya, Chitta Baral, Vivek Gupta 0001
ACL (1)4
2025 GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
abstract
Publicly significant images from events carry valuable contextual information with applications in domains such as journalism and education.However, existing methodologies often struggle to accurately extract this contextual relevance from images.To address this challenge, we introduce GETREASON (Geospatial Event Temporal Reasoning), a framework designed to go beyond surfacelevel image descriptions and infer deeper contextual meaning.We hypothesize that extracting global event, temporal, and geospatial information from an image enables a more accurate understanding of its contextual significance.We also introduce a new metric GREAT (Geospatial, Reasoning and Event Accuracy with Temporal alignment) for a reasoning capturing evaluation.Our layered multi-agentic approach, evaluated using a reasoning-weighted metric, demonstrates that meaningful information can be inferred from images, allowing them to be effectively linked to their corresponding events and broader contextual background.
Shikhhar Siingh, Abhinav Rawat, Chitta Baral, Vivek Gupta 0001
ACL (1)4
2025 Weaver: Interweaving SQL and LLM for Table Reasoning
abstract
Querying tables with unstructured data is challenging due to the presence of text (or image), either embedded in the table or in external paragraphs, which traditional SQL struggles to process, especially for tasks requiring semantic reasoning.While Large Language Models (LLMs) excel at understanding context, they face limitations with long input sequences.Existing approaches that combine SQL and LLM typically rely on rigid, predefined workflows, limiting their adaptability to complex queries.To address these issues, we introduce Weaver , a modular pipeline that dynamically integrates SQL and LLM for table-based question answering (Table QA).Weaver generates a flexible, step-by-step plan that combines SQL for structured data retrieval with LLMs for semantic processing.By decomposing complex queries into manageable subtasks, Weaver improves accuracy and generalization.Our experiments show that Weaver consistently outperforms state-ofthe-art methods across four Table QA datasets, reducing both API calls and error rates.
Rohit Khoja, Devanshu Gupta, Yanjie Fu, Dan Roth 0001, Vivek Gupta 0001
EMNLP5
2025 Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents
abstract
Flowcharts are a critical tool for visualizing decision-making processes.However, their non-linear structure and complex visual-textual relationships make it challenging to interpret them using LLMs, as vision-language models frequently hallucinate nonexistent connections and decision paths when analyzing these diagrams.This leads to compromised reliability for automated flowchart processing in critical domains such as logistics, health, and engineering.We introduce the task of Fine-grained Flowchart Attribution, which traces specific components grounding a flowchart referring LLM response.Flowchart Attribution ensures the verifiability of LLM predictions and improves explainability by linking generated responses to the flowchart's structure.We propose FlowPathAgent, a neurosymbolic agent that performs fine-grained post hoc attribution through graph-based reasoning.It first segments the flowchart, then converts it into a structured symbolic graph, and then employs an agentic approach to dynamically interact with the graph, to generate attribution paths.Additionally, we present FlowExplainBench, a novel benchmark for evaluating flowchart attributions across diverse styles, domains, and question types.Experimental results show that FlowPathAgent mitigates visual hallucinations in LLM answers over flowchart QA, outperforming strong baselines by 10-14% on our proposed FlowExplainBench dataset.
Manan Suri, Puneet Mathur, Nedim Lipka, Franck Dernoncourt, Ryan Rossi, Vivek Gupta 0001, Dinesh Manocha
EMNLP6
2025 H-STAR: LLM-driven Hybrid SQL-Text Adaptive Reasoning on Tables
abstract
Nikhil Abhyankar, Vivek Gupta, Dan Roth, Chandan K. Reddy. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Nikhil Abhyankar, Vivek Gupta 0001, Dan Roth 0001, Chandan K. Reddy
NAACL (Long Papers)2
2025 Leveraging LLM For Synchronizing Information Across Multilingual Tables
abstract
Siddharth Khincha, Tushar Kataria, Ankita Anand, Dan Roth, Vivek Gupta. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Siddharth Khincha, Tushar Kataria, Dan Roth 0001, Vivek Gupta 0001
NAACL (Long Papers)5
2025 MAPWise: Evaluating Vision-Language Models for Advanced Map Queries
abstract
Srija Mukhopadhyay, Abhishek Rajgaria, Prerana Khatiwada, Manish Shrivastava, Dan Roth, Vivek Gupta. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Srija Mukhopadhyay, Abhishek Rajgaria, Prerana Khatiwada, Manish Shrivastava 0001, Dan Roth 0001, Vivek Gupta 0001
NAACL (Long Papers)6
2025 TRANSIENTTABLES: Evaluating LLMs' Reasoning on Temporally Evolving Semi-structured Tables
abstract
Abhilash Shankarampeta, Harsh Mahajan, Tushar Kataria, Dan Roth, Vivek Gupta. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Abhilash Reddy Shankarampeta, Harsh Mahajan, Tushar Kataria, Dan Roth 0001, Vivek Gupta 0001
NAACL (Long Papers)5
2024 Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets
abstract
Language models, characterized by their blackbox nature, often hallucinate and display sensitivity to input perturbations, causing concerns about trust.To enhance trust, it is imperative to gain a comprehensive understanding of the model's failure modes and develop effective strategies to improve their performance.In this study, we introduce a methodology designed to examine how input perturbations affect language models across various scales, including pre-trained models and large language models (LLMs).Utilizing fine-tuning, we enhance the model's robustness to input perturbations.Additionally, we investigate whether exposure to one perturbation enhances or diminishes the model's performance with respect to other perturbations.To address robustness against multiple perturbations, we present three distinct fine-tuning strategies.Furthermore, we broaden the scope of our methodology to encompass large language models (LLMs) by leveraging a chain of thought (CoT) prompting approach augmented with exemplars.We employ the Tabular-NLI task to showcase how our proposed strategies adeptly train a robust model, enabling it to address diverse perturbations while maintaining accuracy on the original dataset.
Pranshu Pandya, Tushar Kataria, Vivek Gupta 0001, Dan Roth 0001
EMNLP4
2023 TempTabQA: Temporal Question Answering for Semi-Structured Tables
abstract
Semi-structured data, such as Infobox tables, often include temporal information about entities, either implicitly or explicitly.Can current NLP systems reason about such information in semi-structured tables?To tackle this question, we introduce the task of temporal question answering on semi-structured tables.We present a dataset, TEMPTABQA, which comprises 11,454 question-answer pairs extracted from 1,208 Wikipedia Infobox tables spanning more than 90 distinct domains.Using this dataset, we evaluate several state-ofthe-art models for temporal reasoning.We observe that even the top-performing LLMs lag behind human performance by more than 13.5 F1 points.Given these results, our dataset has the potential to serve as a challenging benchmark to improve the temporal reasoning capabilities of NLP models.
Vivek Gupta 0001, Pranshu Kandoi, Mahek Bhavesh Vora, Shuo Zhang 0006, Yujie He 0003, Ridho Reinanda, Vivek Srikumar
EMNLP1
2022 Right for the Right Reason: Evidence Extraction for Trustworthy Tabular Reasoning
abstract
Vivek Gupta, Shuo Zhang, Alakananda Vempala, Yujie He, Temma Choji, Vivek Srikumar. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Vivek Gupta 0001, Shuo Zhang 0006, Alakananda Vempala, Yujie He 0003, Temma Choji, Vivek Srikumar
ACL (1)1
2022 IndicXNLI: Evaluating Multilingual Inference for Indian Languages
abstract
While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited.To this end, we introduce INDICXNLI, an NLI dataset for 11 Indic languages.It has been created by high-quality machine translation of the original English XNLI dataset and our analysis attests to the quality of INDICXNLI.By finetuning different pre-trained LMs on this IN-DICXNLI, we analyze various cross-lingual transfer techniques with respect to the impact of the choice of language models, languages, multi-linguality, mix-language input, etc.These experiments provide us with useful insights into the behaviour of pre-trained models for a diverse set of languages.
Divyanshu Aggarwal, Vivek Gupta 0001, Anoop Kunchukuttan
EMNLP2
2022 Bilingual Tabular Inference: A Case Study on Indic Languages
abstract
Chaitanya Agarwal, Vivek Gupta, Anoop Kunchukuttan, Manish Shrivastava. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Chaitanya Agarwal, Vivek Gupta 0001, Anoop Kunchukuttan, Manish Shrivastava 0001
NAACL-HLT2
2022 Is My Model Using The Right Evidence? Systematic Probes for Examining Evidence-Based Tabular Reasoning
abstract
Abstract Neural models command state-of-the-art performance across NLP tasks, including ones involving “reasoning”. Models claiming to reason about the evidence presented to them should attend to the correct parts of the input while avoiding spurious patterns therein, be self-consistent in their predictions across inputs, and be immune to biases derived from their pre-training in a nuanced, context- sensitive fashion. Do the prevalent *BERT- family of models do so? In this paper, we study this question using the problem of reasoning on tabular data. Tabular inputs are especially well-suited for the study—they admit systematic probes targeting the properties listed above. Our experiments demonstrate that a RoBERTa-based model, representative of the current state-of-the-art, fails at reasoning on the following counts: it (a) ignores relevant parts of the evidence, (b) is over- sensitive to annotation artifacts, and (c) relies on the knowledge encoded in the pre-trained language model rather than the evidence presented in its tabular inputs. Finally, through inoculation experiments, we show that fine- tuning the model on perturbed data does not help it overcome the above challenges.
Vivek Gupta 0001, Riyaz A. Bhat, Atreya Ghosal, Manish Shrivastava 0001, Maneesh Kumar Singh 0001, Vivek Srikumar
Trans. Assoc. Comput. Linguistics1
2021 Incorporating External Knowledge to Enhance Tabular Reasoning
abstract
Reasoning about tabular information presents unique challenges to modern NLP approaches which largely rely on pre-trained contextualized embeddings of text.In this paper, we study these challenges through the problem of tabular natural language inference.We propose easy and effective modifications to how information is presented to a model for this task.We show via systematic experiments that these strategies substantially improve tabular inference performance.
J. Neeraja, Vivek Gupta 0001, Vivek Srikumar
NAACL-HLT2
2020 P-SIF: Document Embeddings Using Partition Averaging
abstract
Simple weighted averaging of word vectors often yields effective representations for sentences which outperform sophisticated seq2seq neural models in many tasks. While it is desirable to use the same method to represent documents as well, unfortunately, the effectiveness is lost when representing long documents involving multiple sentences. One of the key reasons is that a longer document is likely to contain words from many different topics; hence, creating a single vector while ignoring all the topical structure is unlikely to yield an effective document representation. This problem is less acute in single sentences and other short text fragments where the presence of a single topic is most likely. To alleviate this problem, we present P-SIF, a partitioned word averaging model to represent long documents. P-SIF retains the simplicity of simple weighted word averaging while taking a document's topical structure into account. In particular, P-SIF learns topic-specific vectors from a document and finally concatenates them all to represent the overall document. We provide theoretical justifications on the correctness of P-SIF. Through a comprehensive set of experiments, we demonstrate P-SIF's effectiveness compared to simple weighted averaging and many other baselines.
Vivek Gupta 0001, Ankit Saw, Pegah Nokhiz, Praneeth Netrapalli, Piyush Rai, Partha P. Talukdar
AAAI1
2020 INFOTABS: Inference on Tables as Semi-structured Data
abstract
In this paper, we observe that semi-structured tabulated text is ubiquitous; understanding them requires not only comprehending the meaning of text fragments, but also implicit relationships between them.We argue that such data can prove as a testing ground for understanding how we reason about information.To study this, we introduce a new dataset called INFOTABS, comprising of human-written textual hypotheses based on premises that are tables extracted from Wikipedia info-boxes.Our analysis shows that the semi-structured, multi-domain and heterogeneous nature of the premises admits complex, multi-faceted reasoning.Experiments reveal that, while human annotators agree on the relationships between a table-hypothesis pair, several standard modeling strategies are unsuccessful at the task, suggesting that reasoning about tables can pose a difficult modeling challenge.
Vivek Gupta 0001, Maitrey Mehta, Pegah Nokhiz, Vivek Srikumar
ACL1
2020 Improving Document Classification with Multi-Sense Embeddings
abstract
Efficient representation of text documents is an important building block in many NLP tasks. Research on long text categorization has shown that simple weighted averaging of word vectors for sentence representation often outperforms more sophisticated neural models. Recently proposed Sparse Composite Document Vector (SCDV) [32] extends this approach from sentences to documents using soft clustering over word vectors. However, SCDV disregards the multi-sense nature of words, and it also suffers from the curse of higher dimensionality. In this work, we address these shortcomings and propose SCDV-MS. SCDV-MS utilizes multi-sense word embeddings and learns a lower dimensional manifold. Through extensive experiments on multiple real-world datasets, we show that SCDV-MS embeddings outperform previous state-of-the-art embeddings on multi-class and multi-label text categorization tasks. Furthermore, SCDV-MS embeddings are more efficient than SCDV in terms of time and space complexity on textual classification tasks. We have released SCDV-MS source code with the paper. (https://github.com/vgupta123/SCDV-MS)
Vivek Gupta 0001, Pegah Nokhiz, Partha P. Talukdar
ECAI1
2019 Distributional Semantics Meets Multi-Label Learning
abstract
We present a label embedding based approach to large-scale multi-label learning, drawing inspiration from ideas rooted in distributional semantics, specifically the Skip Gram Negative Sampling (SGNS) approach, widely used to learn word embeddings. Besides leading to a highly scalable model for multi-label learning, our approach highlights interesting connections between label embedding methods commonly used for multi-label learning and paragraph embedding methods commonly used for learning representations of text data. The framework easily extends to incorporating auxiliary information such as label-label correlations; this is crucial especially when many training instances are only partially annotated. To facilitate end-to-end learning, we develop a joint learning algorithm that can learn the embeddings as well as a regression model that predicts these embeddings for the new input to be annotated, via efficient gradient based methods. We demonstrate the effectiveness of our approach through an extensive set of experiments on a variety of benchmark datasets, and show that the proposed models perform favorably as compared to state-of-the-art methods for large-scale multi-label learning.
Vivek Gupta 0001, Rahul Wadbude, Nagarajan Natarajan, Harish Karnick, Prateek Jain 0002, Piyush Rai
AAAI1
2019 A Logic-Driven Framework for Consistency of Neural Models
abstract
Tao Li, Vivek Gupta, Maitrey Mehta, Vivek Srikumar. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tao Li 0039, Vivek Gupta 0001, Maitrey Mehta, Vivek Srikumar
EMNLP/IJCNLP (1)2
2017 SCDV : Sparse Composite Document Vectors using soft clustering over distributional representations
abstract
We present a feature vector formation technique for documents -Sparse Composite Document Vector (SCDV)which overcomes several shortcomings of the current distributional paragraph vector representations that are widely used for text representation.In SCDV, word embeddings are clustered to capture multiple semantic contexts in which words occur.They are then chained together to form document topic-vectors that can express complex, multi-topic documents.Through extensive experiments on multi-class and multi-label classification tasks, we outperform the previous state-of-the-art method, NTSG (Liu et al., 2015a).We also show that SCDV embeddings perform well on heterogeneous tasks like Topic Coherence, context-sensitive Learning and Information Retrieval.Moreover, we achieve significant reduction in training and prediction times compared to other representation methods.SCDV achieves best of both worlds -better performance with lower time and space complexity.
Dheeraj Mekala, Vivek Gupta 0001, Bhargavi Paranjape, Harish Karnick
EMNLP2
2016 Assisting humans to achieve optimal sleep by changing ambient temperature
abstract
Environment plays a vital role in the sleep mechanism of a human. It has been shown from many studies that sleeping and waking environment, waking time and hours of sleep is of very significant importance [1] which can result in sleeping disorders and variety of diseases. This paper finds the sleep cycle of an individual and according changes the ambient temperature to maximize his/her sleep efficiency. We suggest a method which will assist in increasing sleep efficiency. Using Fast-Fourier-Transformation (FFT) of heart rate signals to extract heart rate variability data such that low frequency / high frequency (LF/HF) power ratio we are detecting sleep stages using an automated algorithm and then applying feedback mechanism to alter the ambient temperature depending upon the sleep stage.
Vivek Gupta 0001, Siddhant Mittal, Sandip Bhaumik, Raj Roy
BIBM1
2016 Product Classification in E-Commerce using Distributional Semantics
abstract
Product classification is the task of automatically predicting a taxonomy path for a product in a predefined taxonomy hierarchy given a textual product description or title. For efficient product classification we require a suitable representation for a document (the textual description of a product) feature vector and efficient and fast algorithms for prediction. To address the above challenges, we propose a new distributional semantics representation for document vector formation. We also develop a new two-level ensemble approach utilising (with respect to the taxonomy tree) path-wise, node-wise and depth-wise classifiers to reduce error in the final product classification task. Our experiments show the effectiveness of the distributional representation and the ensemble approach on data sets from a leading e-commerce platform and achieve improved results on various evaluation metrics compared to earlier approaches.
Vivek Gupta 0001, Harish Karnick, Ashendra Bansal, Pradhuman Jhala
COLING1