VLDB 2026 Research / reviewers in the wild / expert
Tejas Gokhale
dblp:241/9472
· DBLP profile ↗
22ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0002-5593-2804ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Zero-Shot Multimodal Retrieval with Multi-Scale Contextual RepresentationsabstractIn multimodal information retrieval (MMIR), candidates relevant to an input query need to be retrieved from a database, where the query and database items span different modalities.As real-world databases evolve, repeatedly annotating and indexing data and re-optimizing domain-specific models across modalities is impractical.We present Multi-Score, a finetuning-free, two-stage MMIR approach that couples efficient candidate filtering with finegrained multimodal re-ranking.Stage-1 adopts Matryoshka representations to efficiently filter out low-relevance candidates without expensive similarity computations on full-scale representations for the entire database.Stage-2 reranks the filtered candidates by computing their fine-grained multimodal contextual representations with two scoring functions for semantic alignment using chain-of-thought prompting and question-answering.Experiments demonstrate state-of-the-art zero-shot retrieval on 12 MMIR tasks across 32 datasets while outperforming supervised methods on 23 datasets. Sourajit Saha, Tejas Gokhale |
ACL (1) | 2 |
| 2025 | VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical ReasoningabstractMultimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly across multiple images remains a significant challenge. To address this, we introduce VOILA, a large-scale, open-ended, dynamic benchmark designed to evaluate MLLMs' perceptual understanding and abstract relational reasoning. VOILA employs an analogical mapping approach in the visual domain, requiring models to generate an image that completes an analogy between two given image pairs, reference and application, without relying on predefined choices. Our experiments demonstrate that the analogical reasoning tasks in VOILA present a challenge to MLLMs. Through multi-step analysis, we reveal that current MLLMs struggle to comprehend inter-image relationships and exhibit limited capabilities in high-level relational reasoning. Notably, we observe that performance improves when following a multi-step strategy of least-to-most prompting. Comprehensive evaluations on open-source models and GPT-4o show that on text-based answers, the best accuracy for challenging scenarios is 13% (LLaMa 3.2) and even for simpler tasks is only 29% (GPT-4o), while human performance is significantly higher at 70% across both difficulty levels. Nilay Yilmaz, Maitreya Patel, Yiran Luo 0001, Tejas Gokhale, Chitta Baral, Suren Jayasuriya, Yezhou Yang |
ICLR | 4 |
| 2025 | Latent Diffusion Unlearning: Protecting Against Unauthorized Personalization Through Trajectory Shifted PerturbationsabstractText-to-image diffusion models have demonstrated remarkable effectiveness in rapid and high-fidelity personalization, even when provided with only a few user images. However, the effectiveness of personalization techniques has lead to concerns regarding data privacy, intellectual property protection, and unauthorized usage. To mitigate such unauthorized usage and model replication, the idea of generating ''unlearnable'' training samples utilizing image poisoning techniques has emerged. Existing methods for this have limited imperceptibility as they operate in the pixel space which results in images with noise and artifacts. In this work, we propose a novel model-based perturbation strategy that operates within the latent space of diffusion models. Our method alternates between denoising and inversion while modifying the starting point of the denoising trajectory: of diffusion models. This trajectory-shifted sampling ensures that the perturbed images maintain high visual fidelity to the original inputs while being resistant to inversion and personalization by downstream generative models. This approach integrates unlearnability into the framework of Latent Diffusion Models (LDMs), enabling a practical and imperceptible defense against unauthorized model adaptation. We validate our approach on four benchmark datasets to demonstrate robustness against state-of-the-art inversion attacks. Results demonstrate that our method achieves significant improvements in imperceptibility (~8% - 10% on perceptual metrics including PSNR, SSIM, and FID) and robustness (~10% on average across five adversarial settings), highlighting its effectiveness in safeguarding sensitive data. https://github.com/naresh-ub/unlearnable_samples. Naresh Kumar Devulapally, Shruti Agarwal, Tejas Gokhale, Vishnu Suresh Lokhande |
ACM Multimedia | 3 |
| 2025 | Improving Shift Invariance in Convolutional Neural Networks with Translation Invariant Polyphase SamplingabstractConvolutional neural networks (CNNs), widely deployed in several applications, contain downsampling operators in their pooling layers which have been observed to be sensitive to pixel-level shift, affecting the robustness of CNNs. We study shift invariance through the lens of maximum-sampling bias (MSB) and find MSB to be negatively cor-related with shift invariance. Based on this insight, we propose a learnable pooling operator called Translation Invariant Polyphase Sampling (TIPS) to reduce MSB and learn translation-invariant representations. TIPS results in consistent performance gains on multiple benchmarks for image classification, object detection, and semantic segmentation in terms of accuracy, shift consistency, shift fi-delity, as well as improvements in adversarial and distributional robustness. TIPS results in the lowest MSB compared to all previous methods, thus explaining the strong empirical results. TIPS can be integrated into any CNN and can be trained end-to-end with marginal computational overhead. Code: https://github.com/sourajitcs/tips/ Sourajit Saha, Tejas Gokhale |
WACV | 2 |
| 2024 | Towards Robust Visual Understanding: from Recognition to ReasoningabstractModels that learn from data are widely and rapidly being deployed today for real-world use, but they suffer from unforeseen failures due to distribution shift, adversarial attacks, noise and corruption, and data scarcity. But many failures also occur because many modern AI tasks require reasoning beyond pattern matching -- and such reasoning abilities are difficult to formulate as data-based input-output function fitting. The reliability problem has become increasingly important under the new paradigm of semantic ``multimodal'' learning. My research provides avenues to develop robust and reliable computer vision systems, particularly by leveraging the interactions between vision and language. In this AAAI New Faculty highlights talk, I will cover three thematic areas of my research, ranging from robustness in computer vision, open-domain reliability in visual reasoning, and challenges and opportunities in evaluation of generative models. Readers are encouraged to refer to my website (www.tejasgokhale.com) for more details and updates from my lab's activities towards the goal of robust visual understanding. Tejas Gokhale |
AAAI | 1 |
| 2024 | ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion ModelsabstractThe ability to understand visual concepts and replicate and compose these concepts from images is a central goal for computer vision. Recent advances in text-to-image (T2I) models have lead to high definition and realistic image quality generation by learning from large databases of images and their descriptions. However, the evaluation of T2I models has focused on photorealism and limited qualitative measures of visual understanding. To quantify the ability of T2I models in learning and synthesizing novel visual concepts (a.k.a. personalized T2I), we introduce ConceptBed, a large-scale dataset that consists of 284 unique visual concepts, and 33K composite text prompts. Along with the dataset, we propose an evaluation metric, Concept Confidence Deviation (CCD), that uses the confidence of oracle concept classifiers to measure the alignment between concepts generated by T2I generators and concepts contained in target images. We evaluate visual concepts that are either objects, attributes, or styles, and also evaluate four dimensions of compositionality: counting, attributes, relations, and actions. Our human study shows that CCD is highly correlated with human understanding of concepts. Our results point to a trade-off between learning the concepts and preserving the compositionality which existing approaches struggle to overcome. The data, code, and interactive demo is available at: https://conceptbed.github.io/ Maitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou Yang |
AAAI | 2 |
| 2024 | On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth EstimationabstractRecent advances in monocular depth estimation have been made by incorporating natural language as additional guidance. Although yielding impressive results, the impact of the language prior, particularly in terms of generalization and robustness, remains unexplored. In this paper, we address this gap by quantifying the impact of this prior and introduce methods to benchmark its effectiveness across various settings. We generate “low-level” sentences that convey object-centric, three-dimensional spatial relationships, incorporate them as additional language priors and evaluate their downstream impact on depth estimation. Our key finding is that current language-guided depth estimators perform optimally only with scene-level descriptions and counter-intuitively fare worse with low level descriptions. Despite leveraging additional data, these methods are not robust to directed adversarial attacks and decline in performance with an increase in distribution shift. Finally, to provide a foundation for future research, we identify points of failures and offer insights to better understand these shortcomings. With an increasing number of methods using language for depth estimation, our findings highlight the opportunities and pitfalls that require careful consideration for effective deployment in real-world settings.11Code/Data: https://github.com/agneet42/lang_depth Agneet Chatterjee, Tejas Gokhale, Chitta Baral, Yezhou Yang |
CVPR | 2 |
| 2024 | REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models
Agneet Chatterjee, Yiran Luo 0001, Tejas Gokhale, Yezhou Yang, Chitta Baral |
ECCV (30) | 3 |
| 2024 | Getting it Right: Improving Spatial Consistency in Text-to-Image Models
Agneet Chatterjee, Gabriela Ben Melech Stan, Estelle Aflalo, Sayak Paul, Dhruba Ghosh, Tejas Gokhale, Ludwig Schmidt, Hannaneh Hajishirzi, Vasudev Lal, Chitta Baral, Yezhou Yang |
ECCV (22) | 6 |
| 2024 | TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language NegativesabstractContrastive Language-Image Pretraining (CLIP) models maximize the mutual information between text and visual modalities to learn representations. This makes the nature of the training data a significant factor in the efficacy of CLIP for downstream tasks. However, the lack of compositional diversity in contemporary image-text datasets limits the compositional reasoning ability of CLIP. We show that generating ``hard'' negative captions via in-context learning and synthesizing corresponding negative images with text-to-image generators offers a solution. We introduce a novel contrastive pre-training strategy that leverages these hard negative captions and images in an alternating fashion to train CLIP. We demonstrate that our method, named TripletCLIP, when applied to existing datasets such as CC3M and CC12M, enhances the compositional capabilities of CLIP, resulting in an absolute improvement of over 9% on the SugarCrepe benchmark on an equal computational budget, as well as improvements in zero-shot image classification and image retrieval. Our code, models, and data are available at: tripletclip.github.io. Maitreya Patel, Abhiram Kusumba, Changhoon Kim, Tejas Gokhale, Chitta Baral, Yezhou Yang |
NeurIPS | 5 |
| 2023 | End-to-end Knowledge Retrieval with Multi-modal QueriesabstractWe investigate knowledge retrieval with multimodal queries, i.e. queries containing information split across image and text inputs, a challenging task that differs from previous work on cross-modal retrieval.We curate a new dataset called ReMuQ 1 for benchmarking progress on this task.ReMuQ requires a system to retrieve knowledge from a large corpus by integrating contents from both text and image queries.We introduce a retriever model "ReViz" that can directly process input text and images to retrieve relevant knowledge in an end-to-end fashion without being dependent on intermediate modules such as object detectors or caption generators.We introduce a new pretraining task that is effective for learning knowledge retrieval with multimodal queries and also improves performance on downstream tasks.We demonstrate superior performance in retrieval on two datasets (ReMuQ and OK-VQA) under zeroshot settings as well as further improvements when finetuned on these datasets. Man Luo 0003, Zhiyuan Fang, Tejas Gokhale, Yezhou Yang, Chitta Baral |
ACL (1) | 3 |
| 2023 | Adversarial Bayesian Augmentation for Single-Source Domain GeneralizationabstractGeneralizing to unseen image domains is a challenging problem primarily due to the lack of diverse training data, inaccessible target data, and the large domain shift that may exist in many real-world settings. As such data augmentation is a critical component of domain generalization methods that seek to address this problem. We present Adversarial Bayesian Augmentation (ABA), a novel algorithm that learns to generate image augmentations in the challenging single-source domain generalization setting. ABA draws on the strengths of adversarial learning and Bayesian neural networks to guide the generation of diverse data augmentations –these synthesized image domains aid the classifier in generalizing to unseen domains. We demonstrate the strength of ABA on several types of domain shift including style shift, subpopulation shift, and shift in the medical imaging setting. ABA outperforms all previous state-of-the-art methods, including pre-specified augmentations, pixel-based and convolutional-based augmentations. Code: https://github.com/shengcheng/ABA. Tejas Gokhale, Yezhou Yang |
ICCV | 2 |
| 2023 | Improving Diversity with Adversarially Learned Transformations for Domain GeneralizationabstractTo be successful in single source domain generalization (SSDG), maximizing diversity of synthesized domains has emerged as one of the most effective strategies. Recent success in SSDG comes from methods that pre-specify diversity inducing image augmentations during training, so that it may lead to better generalization on new domains. However, naïve pre-specified augmentations are not always effective, either because they cannot model large domain shift, or be-cause the specific choice of transforms may not cover the types of shift commonly occurring in domain generalization. To address this issue, we present a novel framework called ALT: adversarially learned transformations, that uses an adversary neural network to model plausible, yet hard image transformations that fool the classifier. ALT learns image transformations by randomly initializing the adversary net-work for each batch and optimizing it for a fixed number of steps to maximize classification error. The classifier is trained by enforcing a consistency between its predictions on the clean and transformed images. With extensive empirical analysis, we find that this new form of adversarial transformations achieves both objectives of diversity and hardness simultaneously, outperforming all existing techniques on competitive benchmarks for SSDG. We also show that ALT can seamlessly work with existing diversity modules to produce highly distinct, and large transformations of the source domain leading to state-of-the-art performance. Code: https://github.com/tejas-gokhale/ALT Tejas Gokhale, Rushil Anirudh, Jayaraman J. Thiagarajan, Bhavya Kailkhura, Chitta Baral, Yezhou Yang |
WACV | 1 |
| 2022 | Improving Biomedical Information Retrieval with Neural RetrieversabstractInformation retrieval (IR) is essential in search engines and dialogue systems as well as natural language processing tasks such as open-domain question answering. IR serve an important function in the biomedical domain, where content and sources of scientific knowledge may evolve rapidly. Although neural retrievers have surpassed traditional IR approaches such as TF-IDF and BM25 in standard open-domain question answering tasks, they are still found lacking in the biomedical domain. In this paper, we seek to improve information retrieval (IR) using neural retrievers (NR) in the biomedical domain, and achieve this goal using a three-pronged approach. First, to tackle the relative lack of data in the biomedical domain, we propose a template-based question generation method that can be leveraged to train neural retriever models. Second, we develop two novel pre-training tasks that are closely aligned to the downstream task of information retrieval. Third, we introduce the ``Poly-DPR'' model which encodes each context into multiple context vectors. Extensive experiments and analysis on the BioASQ challenge suggest that our proposed method leads to large gains over existing neural approaches and beats BM25 in the small-corpus setting. We show that BM25 and our method can complement each other, and a simple hybrid model leads to further gains in the large corpus setting. Man Luo 0003, Arindam Mitra, Tejas Gokhale, Chitta Baral |
AAAI | 3 |
| 2022 | CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question AnsweringabstractVideos often capture objects, their visible properties, their motion, and the interactions between different objects.Objects also have physical properties such as mass, which the imaging pipeline is unable to directly capture.However, these properties can be estimated by utilizing cues from relative object motion and the dynamics introduced by collisions.In this paper, we introduce CRIPP-VQA 1 , a new video question answering dataset for reasoning about the implicit physical properties of objects in a scene.CRIPP-VQA contains videos of objects in motion, annotated with questions that involve counterfactual reasoning about the effect of actions, questions about planning in order to reach a goal, and descriptive questions about visible properties of objects.The CRIPP-VQA test set enables evaluation under several outof-distribution settings -videos with objects with masses, coefficients of friction, and initial velocities that are not observed in the training distribution.Our experiments reveal a surprising and significant performance gap in terms of answering questions about implicit properties (the focus of this paper) and explicit properties of objects (the focus of prior work). Maitreya Patel, Tejas Gokhale, Chitta Baral, Yezhou Yang |
EMNLP | 2 |
| 2021 | Attribute-Guided Adversarial Training for Robustness to Natural PerturbationsabstractWhile existing work in robust deep learning has focused on small pixel-level norm-based perturbations, this may not account for perturbations encountered in several real world settings. In many such cases although test data might not be available, broad specifications about the types of perturbations (such as an unknown degree of rotation) may be known. We consider a setup where robustness is expected over an unseen test domain that is not i.i.d. but deviates from the training domain. While this deviation may not be exactly known, its broad characterization is specified a priori, in terms of attributes. We propose an adversarial training approach which learns to generate new samples so as to maximize exposure of the classifier to the attributes-space, without having access to the data from the test domain. Our adversarial training solves a min-max optimization problem, with the inner maximization generating adversarial perturbations, and the outer minimization finding model parameters by optimizing the loss on adversarial perturbations generated from the inner maximization. We demonstrate the applicability of our approach on three types of naturally occurring perturbations --- object-related shifts, geometric transformations, and common image corruptions. Our approach enables deep neural networks to be robust against a wide range of naturally occurring perturbations. We demonstrate the usefulness of the proposed approach by showing the robustness gains of deep neural networks trained using our adversarial training on MNIST, CIFAR-10, and a new variant of the CLEVR dataset. Tejas Gokhale, Rushil Anirudh, Bhavya Kailkhura, Jayaraman J. Thiagarajan, Chitta Baral, Yezhou Yang |
AAAI | 1 |
| 2021 | Weakly Supervised Relative Spatial Reasoning for Visual Question AnsweringabstractVision-and-language (V&L) reasoning necessitates perception of visual concepts such as objects and actions, understanding semantics and language grounding, and reasoning about the interplay between the two modalities. One crucial aspect of visual reasoning is spatial understanding, which involves understanding relative locations of objects, i.e. implicitly learning the geometry of the scene. In this work, we evaluate the faithfulness of V&L models to such geometric understanding, by formulating the prediction of pair-wise relative locations of objects as a classification as well as a regression task. Our findings suggest that state-of-the-art transformer-based V&L models lack sufficient abilities to excel at this task. Motivated by this, we design two objectives as proxies for 3D spatial reasoning (SR) – object centroid estimation, and relative position estimation, and train V&L with weak supervision from off-the-shelf depth estimators. This leads to considerable improvements in accuracy for the "GQA" visual question answering challenge (in fully supervised, few-shot, and O.O.D settings) as well as improvements in relative spatial reasoning. Code and data will be released here. Pratyay Banerjee, Tejas Gokhale, Yezhou Yang, Chitta Baral |
ICCV | 2 |
| 2021 | Self-Supervised Test-Time Learning for Reading ComprehensionabstractRecent work on unsupervised question answering has shown that models can be trained with procedurally generated question-answer pairs and can achieve performance competitive with supervised methods.In this work, we consider the task of unsupervised reading comprehension and present a method that performs "test-time learning" (TTL) on a given context (text passage), without requiring training on large-scale human-authored datasets containing context-question-answer triplets.This method operates directly on a single test context, uses self-supervision to train models on synthetically generated question-answer pairs, and then infers answers to unseen humanauthored questions for this context.Our method achieves accuracies competitive with fully supervised methods and significantly outperforms current unsupervised methods.TTL methods with a smaller model are also competitive with the current state-of-the-art in unsupervised reading comprehension. Pratyay Banerjee, Tejas Gokhale, Chitta Baral |
NAACL-HLT | 2 |
| 2020 | VQA-LOL: Visual Question Answering Under the Lens of Logic
Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang |
ECCV (21) | 1 |
| 2020 | Video2Commonsense: Generating Commonsense Descriptions to Enrich Video CaptioningabstractCaptioning is a crucial and challenging task for video understanding.In videos that involve active agents such as humans, the agent's actions can bring about myriad changes in the scene.Observable changes such as movements, manipulations, and transformations of the objects in the scene, are reflected in conventional video captioning.Unlike images, actions in videos are also inherently linked to social aspects such as intentions (why the action is taking place), effects (what changes due to the action), and attributes that describe the agent.Thus for video understanding, such as when captioning videos or when answering questions about videos, one must have an understanding of these commonsense aspects.We present the first work on generating commonsense captions directly from videos, to describe latent aspects such as intentions, effects, and attributes.We present a new dataset "Video-to-Commonsense (V2C)" that contains ∼ 9k videos of human agents performing various actions, annotated with 3 types of commonsense descriptions.Additionally we explore the use of open-ended video-based commonsense question answering (V2C-QA) as a way to enrich our captions.Both the generation task and the QA task can be used to enrich video captions.. frame frame frame CNN Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang |
EMNLP (1) | 2 |
| 2020 | MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question AnsweringabstractWhile progress has been made on the visual question answering leaderboards, models often utilize spurious correlations and priors in datasets under the i.i.d.setting.As such, evaluation on out-of-distribution (OOD) test samples has emerged as a proxy for generalization.In this paper, we present MUTANT, a training paradigm that exposes the model to perceptually similar, yet semantically distinct mutations of the input, to improve OOD generalization, such as the VQA-CP challenge.Under this paradigm, models utilize a consistency-constrained training objective to understand the effect of semantic changes in input (question-image pair) on the output (answer).Unlike existing methods on VQA-CP, MUTANT does not rely on the knowledge about the nature of train and test answer distributions.MUTANT establishes a new state-ofthe-art accuracy on VQA-CP with a 10.57% improvement.Our work opens up avenues for the use of semantic input mutations for OOD generalization in question answering. Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang |
EMNLP (1) | 1 |
| 2019 | Vision beyond Pixels: Visual Reasoning via Blocksworld AbstractionsabstractDeep neural networks trained in an end-to-end fashion have brought about exceptional advances in computer vision, especially in computational perception. We go beyond perception and seek to enable vision modules to reason about perceived visual entities such as scenes, objects and actions. We introduce a challenging visual reasoning task, Image-Based Event Sequencing (IES) and compile the first IES dataset, Blocksworld Image Reasoning Dataset (BIRD). Motivated by the blocksworld concept, we propose a modular approach supported by literature in cognitive psychology and children's development. We decompose the problem into two stages - visual perception and event sequencing, and show that our approach can be extended to natural images without re-training. Tejas Gokhale |
IJCAI | 1 |