Arjun R. Akula

dblp:152/3930 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 8 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Visual Intention Grounding for Egocentric Assistants
Pengzhan Sun 0001, Junbin Xiao, Tze Ho Elden Tse, Yicong Li 0004, Arjun R. Akula, Angela Yao
ICCV5
2023 MetaCLUE: Towards Comprehensive Visual Metaphors Research
abstract
Creativity is an indispensable part of human cognition and also an inherent part of how we make sense of the world. Metaphorical abstraction is fundamental in communicating creative ideas through nuanced relationships between abstract concepts such as feelings. While computer vision benchmarks and approaches predominantly focus on understanding and generating literal interpretations of images, metaphorical comprehension of images remains relatively unexplored. Towards this goal, we introduce Meta-CLUE, a set of vision tasks on visual metaphor. We also collect high-quality and rich metaphor annotations (abstract objects, concepts, relationships along with their corresponding object boxes) as there do not exist any datasets that facilitate the evaluation of these tasks. We perform a comprehensive analysis of state-of-the-art models in vision and language based on our annotations, highlighting strengths and weaknesses of current approaches in visual metaphor classification, localization, understanding (retrieval, question answering, captioning) and generation (text-to-image synthesis) tasks. We hope this work provides a concrete step towards developing AI systems with human-like creative capabilities. Project page: https://metaclue.github.io
Arjun R. Akula, Brendan Driscoll, Pradyumna Narayana, Soravit Changpinyo, Zhiwei Jia, Suyash Damle, Garima Pruthi, Sugato Basu, Leonidas J. Guibas, William T. Freeman, Yuanzhen Li, Varun Jampani
CVPR1
2023 Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
Weixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani, Arjun R. Akula, Pradyumna Narayana, Sugato Basu, Xin Wang 0061, William Yang Wang
ICLR5
2023 LayoutGPT: Compositional Visual Planning and Generation with Large Language Models
abstract
Attaining a high degree of user controllability in visual generation often requires intricate, fine-grained inputs like layouts. However, such inputs impose a substantial burden on users when compared to simple text inputs. To address the issue, we study how Large Language Models (LLMs) can serve as visual planners by generating layouts from text conditions, and thus collaborate with visual generative models. We propose LayoutGPT, a method to compose in-context visual demonstrations in style sheet language to enhance visual planning skills of LLMs. We show that LayoutGPT can generate plausible layouts in multiple domains, ranging from 2D images to 3D indoor scenes. LayoutGPT also shows superior performance in converting challenging language concepts like numerical and spatial relations to layout arrangements for faithful text-to-image generation. When combined with a downstream image generation model, LayoutGPT outperforms text-to-image models/systems by 20-40\% and achieves comparable performance as human users in designing visual layouts for numerical and spatial correctness. Lastly, LayoutGPT achieves comparable performance to supervised methods in 3D indoor scene synthesis, demonstrating its effectiveness and potential in multiple visual domains.
Weixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani, Arjun R. Akula, Xuehai He, Sugato Basu, Xin Wang 0061, William Yang Wang
NeurIPS5
2022 ALFRED-L: Investigating the Role of Language for Action Learning in Interactive Visual Environments
abstract
Arjun Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tur. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Arjun R. Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tür
EMNLP1
2022 CPL: Counterfactual Prompt Learning for Vision and Language Models
abstract
Xuehai He, Diji Yang, Weixi Feng, Tsu-Jui Fu, Arjun Akula, Varun Jampani, Pradyumna Narayana, Sugato Basu, William Yang Wang, Xin Wang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Xuehai He, Diji Yang, Weixi Feng, Tsu-Jui Fu, Arjun R. Akula, Varun Jampani, Pradyumna Narayana, Sugato Basu, William Yang Wang, Xin Wang 0061
EMNLP5
2021 CrossVQA: Scalably Generating Benchmarks for Systematically Testing VQA Generalization
abstract
One challenge in evaluating visual question answering (VQA) models in the cross-dataset adaptation setting is that the distribution shifts are multi-modal, making it difficult to identify if it is the shifts in visual or language features that play a key role.In this paper, we propose a semi-automatic framework for generating disentangled shifts by introducing a controllable visual question-answer generation (VQAG) module that is capable of generating highly-relevant and diverse questionanswer pairs with the desired dataset style.We use it to create CrossVQA, a collection of test splits for assessing VQA generalization based on the VQA2, VizWiz, and Open Images datasets.We provide an analysis of our generated datasets and demonstrate its utility by using them to evaluate several state-of-theart VQA systems.One important finding is that the visual shifts in cross-dataset VQA matter more than the language shifts.More broadly, we present a scalable framework for systematically evaluating the machine with little human intervention.
Arjun R. Akula, Soravit Changpinyo, Boqing Gong, Piyush Sharma, Song-Chun Zhu, Radu Soricut
EMNLP (1)1
2021 Mind the Context: The Impact of Contextualization in Neural Module Networks for Grounding Visual Referring Expressions
abstract
Neural module networks (NMN) are a popular approach for grounding visual referring expressions.Prior implementations of NMN use pre-defined and fixed textual inputs in their module instantiation.This necessitates a large number of modules as they lack the ability to share weights and exploit associations between similar textual contexts (e.g."dark cube on the left" vs. "black cube on the left").In this work, we address these limitations and evaluate the impact of contextual clues in improving the performance of NMN models.First, we address the problem of fixed textual inputs by parameterizing the module arguments.This substantially reduce the number of modules in NMN by up to 75% without any loss in performance.Next we propose a method to contextualize our parameterized model to enhance the module's capacity in exploiting the visiolinguistic associations.Our model outperforms the state-of-the-art NMN model on CLEVR-Ref+ dataset with +8.1% improvement in accuracy on the single-referent test set and +4.3% on the full test set.Additionally, we demonstrate that contextualization provides +11.2% and +1.7% improvements in accuracy over prior NMN models on CLO-SURE and NLVR2.We further evaluate the impact of our contextualization by constructing a contrast set for CLEVR-Ref+, which we call CC-Ref+.We significantly outperform the baselines by as much as +10.4% absolute accuracy on CC-Ref+, illustrating the generalization skills of our approach.Our dataset is publicly available at https://github.com/ McGill-NLP/contextual-nmn.
Arjun R. Akula, Spandana Gella, Keze Wang, Song-Chun Zhu, Siva Reddy
EMNLP (1)1
2021 Robust Visual Reasoning via Language Guided Neural Module Networks
abstract
Neural module networks (NMN) are a popular approach for solving multi-modal tasks such as visual question answering (VQA) and visual referring expression recognition (REF). A key limitation in prior implementations of NMN is that the neural modules do not effectively capture the association between the visual input and the relevant neighbourhood context of the textual input. This limits their generalizability. For instance, NMN fail to understand new concepts such as “yellow sphere to the left" even when it is a combination of known concepts from train data: “blue sphere", “yellow cube", and “metallic cube to the left". In this paper, we address this limitation by introducing a language-guided adaptive convolution layer (LG-Conv) into NMN, in which the filter weights of convolutions are explicitly multiplied with a spatially varying language-guided kernel. Our model allows the neural module to adaptively co-attend over potential objects of interest from the visual and textual inputs. Extensive experiments on VQA and REF tasks demonstrate the effectiveness of our approach. Additionally, we propose a new challenging out-of-distribution test split for REF task, which we call C3-Ref+, for explicitly evaluating the NMN’s ability to generalize well to adversarial perturbations and unseen combinations of known concepts. Experiments on C3-Ref+ further demonstrate the generalization capabilities of our approach.
Arjun R. Akula, Varun Jampani, Soravit Changpinyo, Song-Chun Zhu
NeurIPS1
2020 CoCoX: Generating Conceptual and Counterfactual Explanations via Fault-Lines
abstract
We present CoCoX (short for Conceptual and Counterfactual Explanations), a model for explaining decisions made by a deep convolutional neural network (CNN). In Cognitive Psychology, the factors (or semantic-level features) that humans zoom in on when they imagine an alternative to a model prediction are often referred to as fault-lines. Motivated by this, our CoCoX model explains decisions made by a CNN using fault-lines. Specifically, given an input image I for which a CNN classification model M predicts class cpred, our fault-line based explanation identifies the minimal semantic-level features (e.g., stripes on zebra, pointed ears of dog), referred to as explainable concepts, that need to be added to or deleted from I in order to alter the classification category of I by M to another specified class calt. We argue that, due to the conceptual and counterfactual nature of fault-lines, our CoCoX explanations are practical and more natural for both expert and non-expert users to understand the internal workings of complex deep learning models. Extensive quantitative and qualitative experiments verify our hypotheses, showing that CoCoX significantly outperforms the state-of-the-art explainable AI models. Our implementation is available at https://github.com/arjunakula/CoCoX
Arjun R. Akula, Song-Chun Zhu
AAAI1
2020 Words Aren't Enough, Their Order Matters: On the Robustness of Grounding Visual Referring Expressions
abstract
Visual referring expression recognition is a challenging task that requires natural language understanding in the context of an image.We critically examine RefCOCOg, a standard benchmark for this task, using a human study and show that 83.7% of test instances do not require reasoning on linguistic structure, i.e., words are enough to identify the target object, the word order doesn't matter.To measure the true progress of existing models, we split the test set into two sets, one which requires reasoning on linguistic structure and the other which doesn't.Additionally, we create an out-of-distribution dataset Ref-Adv by asking crowdworkers to perturb in-domain examples such that the target object changes.Using these datasets, we empirically show that existing methods fail to exploit linguistic structure and are 12% to 23% lower in performance than the established progress for this task.We also propose two methods, one based on contrastive learning and the other based on multi-task learning, to increase the robustness of ViLBERT, the current state-ofthe-art model for this task.Our datasets are publicly
Arjun R. Akula, Spandana Gella, Yaser Al-Onaizan, Song-Chun Zhu, Siva Reddy
ACL1
2014 Towards Auto-remediation in Services Delivery: Context-Based Classification of Noisy and Unstructured Tickets
Gargi Dasgupta, Tapan Kumar Nayak, Arjun R. Akula, Shivali Agarwal, Shripad Nadgowda
ICSOC3
2013 A Novel Approach Towards Incorporating Context Processing Capabilities in NLIDB System
Arjun R. Akula, Rajeev Sangal, Radhika Mamidi
IJCNLP1