Spandana Gella

dblp:146/3968 · DBLP profile ↗
← Back
24ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0003-2725-4476ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Augmenting LLM Reasoning with Dynamic Notes Writing for Complex MultiHop QA
Rishabh Maheshwary, Masoud Hashemi, Khyati Mahajan, Shiva Krishna Reddy Malay, Sai Rajeswar, Sathwik Tejaswi Madhusudhan, Spandana Gella, Vikas Yadav
LREC7
2025 WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
abstract
Rabiul Awal, Mahsa Massoud, Aarash Feizi, Zichao Li, Suyuchen Wang, Christopher Pal, Aishwarya Agrawal, David Vazquez, Siva Reddy, Juan A. Rodriguez, Perouz Taslakian, Spandana Gella, Sai Rajeswar. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Rabiul Awal, Mahsa Massoud, Aarash Feizi, Suyuchen Wang, Christopher Joseph Pal, Aishwarya Agrawal, David Vázquez 0001, Siva Reddy, Juan A. Rodríguez, Perouz Taslakian, Spandana Gella, Sai Rajeswar
EMNLP12
2025 UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
abstract
Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments, critical for many professional and everyday tasks, remain underexplored due to data collection challenges and licensing issues. We introduce UI-Vision, the first comprehensive, license-permissive benchmark for offline, fine-grained evaluation of computer use agents in real-world desktop environments. Unlike online benchmarks, UI-Vision provides: (i) dense, high-quality annotations of human demonstrations, including bounding boxes, UI labels, and action trajectories (clicks, drags, and keyboard inputs) across 83 software applications, and (ii) three fine-to-coarse grained tasks—Element Grounding, Layout Grounding, and Action Prediction—with well-defined metrics to rigorously evaluate agents’ performance in desktop environments. Our evaluation reveals critical limitations in state-of-the-art models like UI-TARS-72B, including issues with understanding professional software, spatial reasoning, and complex actions like drag-and-drop. These findings highlight the challenges in developing fully autonomous computer-use agents. With UI-Vision, we aim to advance the development of more capable agents for real-world desktop tasks.
Shravan Nayak, Xiangru Jian, Qinghong Lin, Juan A. Rodríguez, Montek Kalsi, Nicolas Chapados, M. Tamer Özsu, Aishwarya Agrawal, David Vázquez 0001, Christopher Joseph Pal, Perouz Taslakian, Spandana Gella, Sai Rajeswar
ICML12
2025 SafeArena: Evaluating the Safety of Autonomous Web Agents
abstract
LLM-based agents are becoming increasingly proficient at solving web-based tasks. With this capability comes a greater risk of misuse for malicious purposes, such as posting misinformation in an online forum or selling illicit substances on a website. To evaluate these risks, we propose SafeArena, a benchmark focused on the deliberate misuse of web agents. SafeArena comprises 250 safe and 250 harmful tasks across four websites. We classify the harmful tasks into five harm categories—misinformation, illegal activity, harassment, cybercrime, and social bias, designed to assess realistic misuses of web agents. We evaluate leading LLM-based web agents, including GPT-4o, Claude-3.5 Sonnet, Qwen-2-VL 72B, and Llama-3.2 90B, on our benchmark. To systematically assess their susceptibility to harmful tasks, we introduce the Agent Risk Assessment framework that categorizes agent behavior across four risk levels. We find agents are surprisingly compliant with malicious requests, with GPT-4o and Qwen-2 completing 34.7% and 27.3% of harmful requests, respectively. Our findings highlight the urgent need for safety alignment procedures for web agents.
Ada Defne Tur, Nicholas Meade, Xing Han Lù, Alejandra Zambrano, Arkil Patel, Esin Durmus, Spandana Gella, Karolina Stanczak, Siva Reddy
ICML7
2025 AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document Understanding
abstract
Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps visual features generated by a vision encoder to a shared embedding space with the LLM while preserving semantic similarity. Existing connectors, such as multilayer perceptrons (MLPs), lack inductive bias to constrain visual features within the linguistic structure of the LLM’s embedding space, making them data-hungry and prone to cross-modal misalignment. In this work, we propose a novel vision-text alignment method, AlignVLM, that maps visual features to a weighted average of LLM text embeddings. Our approach leverages the linguistic priors encoded by the LLM to ensure that visual features are mapped to regions of the space that the LLM can effectively interpret. AlignVLM is particularly effective for document understanding tasks, where visual and textual modalities are highly correlated. Our extensive experiments show that AlignVLM achieves state-of-the-art performance compared to prior alignment methods, with larger gains on document understanding and under low-resource setups. We provide further analysis demonstrating its efficiency and robustness to noise.
Ahmed Masry, Juan A. Rodríguez, Suyuchen Wang, Aarash Feizi, Akshay Kalkunte Suresh, Abhay Puri, Xiangru Jian, Pierre-André Noël, Sathwik Tejaswi Madhusudhan, Marco Pedersoli, Bang Liu 0003, Nicolas Chapados, Yoshua Bengio, Enamul Hoque Prince, Christopher Joseph Pal, Issam H. Laradji, David Vázquez 0001, Perouz Taslakian, Spandana Gella, Sai Rajeswar
NeurIPS21
2025 Rendering-Aware Reinforcement Learning for Vector Graphics Generation
abstract
Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-quality SVG generation by framing the problem as a code generation task and leveraging large-scale pretraining. VLMs are particularly suitable for this task as they capture both global semantics and fine-grained visual patterns, while transferring knowledge across vision, natural language, and code domains. However, existing VLM approaches often struggle to produce faithful and efficient SVGs because they never observe the rendered images during training. Although differentiable rendering for autoregressive SVG code generation remains unavailable, rendered outputs can still be compared to original inputs, enabling evaluative feedback suitable for reinforcement learning (RL). We introduce Reinforcement Learning from Rendering Feedback, an RL method that enhances SVG generation in autoregressive VLMs by leveraging feedback from rendered SVG outputs. Given an input image, the model generates SVG roll-outs that are rendered and compared to the original image to compute a reward. This visual fidelity feedback guides the model toward producing more accurate, efficient, and semantically coherent SVGs. \method significantly outperforms supervised fine-tuning, addressing common failure modes and enabling precise, high-quality SVG generation with strong structural understanding and generalization.
Juan A. Rodríguez, Abhay Puri, Rishav Pramanik, Aarash Feizi, Pascal Wichmann, Arnab Kumar Mondal, Mohammad Reza Samsami, Rabiul Awal, Perouz Taslakian, Spandana Gella, Sai Rajeswar, David Vázquez 0001, Christopher Joseph Pal, Marco Pedersoli
NeurIPS11
2023 Multimodal Embodied Plan Prediction Augmented with Synthetic Embodied Dialogue
abstract
Embodied task completion is a challenge where an agent in a simulated environment must predict environment actions to complete tasks based on natural language instructions and egocentric visual observations.We propose a variant of this problem where the agent predicts actions at a higher level of abstraction called a plan, which helps make agent actions more interpretable and can be obtained from the appropriate prompting of large language models.We show that multimodal transformer models can outperform language-only models for this problem but fall significantly short of oracle plans.Since collecting human-human dialogues for embodied environments is expensive and time-consuming, we propose a method to synthetically generate such dialogues, which we then use as training data for plan prediction.We demonstrate that multimodal transformer models can attain strong zero-shot performance from our synthetic data, outperforming language-only models trained on humanhuman data.
Aishwarya Padmakumar, Mert Inan, Spandana Gella, Patrick Lange, Dilek Hakkani-Tür
EMNLP3
2023 "What do others think?": Task-Oriented Conversational Modeling with Subjective Knowledge
abstract
Chao Zhao, Spandana Gella, Seokhwan Kim, Di Jin, Devamanyu Hazarika, Alexandros Papangelis, Behnam Hedayatnia, Mahdi Namazifar, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 24th Meeting of the Special Interest Group on Discourse and Dialogue. 2023.
Spandana Gella, Seokhwan Kim, Di Jin 0005, Devamanyu Hazarika, Alexandros Papangelis, Behnam Hedayatnia, Mahdi Namazifar, Yang Liu 0004, Dilek Hakkani-Tür
SIGDIAL2
2022 TEACh: Task-Driven Embodied Agents That Chat
abstract
Robots operating in human spaces must be able to engage in natural language interaction, both understanding and executing instructions, and using conversation to resolve ambiguity and correct mistakes. To study this, we introduce TEACh, a dataset of over 3,000 human-human, interactive dialogues to complete household tasks in simulation. A Commander with access to oracle information about a task communicates in natural language with a Follower. The Follower navigates through and interacts with the environment to complete tasks varying in complexity from "Make Coffee" to "Prepare Breakfast", asking questions and getting additional information from the Commander. We propose three benchmarks using TEACh to study embodied intelligence challenges, and we evaluate initial models' abilities in dialogue understanding, language grounding, and task execution.
Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gökhan Tür, Dilek Hakkani-Tür
AAAI6
2022 ALFRED-L: Investigating the Role of Language for Action Learning in Interactive Visual Environments
abstract
Arjun Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tur. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Arjun R. Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tür
EMNLP2
2022 Dialog Acts for Task Driven Embodied Agents
abstract
Embodied agents need to be able to interact in natural language -understanding task descriptions and asking appropriate follow up questions to obtain necessary information to be effective at successfully accomplishing tasks for a wide range of users.In this work, we propose a set of dialog acts for modelling such dialogs and annotate the TEACh dataset that includes over 3,000 situated, task oriented conversations (consisting of 39.5k utterances in total) with dialog acts.TEACh-DA is one of the first large scale dataset of dialog act annotations for embodied task completion.Furthermore, we demonstrate the use of this annotated dataset in training models for tagging the dialog acts of a given utterance, predicting the dialog act of the next response given a dialog history, and use the dialog acts to guide agent's non-dialog behaviour.In particular, our experiments on the TEACh Execution from Dialog History task where the model predicts the sequence of low level actions to be executed in the environment for embodied task completion, demonstrate that dialog acts can improve end task success rate by up to 2 points compared to the system without dialog acts.
Spandana Gella, Aishwarya Padmakumar, Patrick Lange, Dilek Hakkani-Tür
SIGDIAL1
2021 Mind the Context: The Impact of Contextualization in Neural Module Networks for Grounding Visual Referring Expressions
abstract
Neural module networks (NMN) are a popular approach for grounding visual referring expressions.Prior implementations of NMN use pre-defined and fixed textual inputs in their module instantiation.This necessitates a large number of modules as they lack the ability to share weights and exploit associations between similar textual contexts (e.g."dark cube on the left" vs. "black cube on the left").In this work, we address these limitations and evaluate the impact of contextual clues in improving the performance of NMN models.First, we address the problem of fixed textual inputs by parameterizing the module arguments.This substantially reduce the number of modules in NMN by up to 75% without any loss in performance.Next we propose a method to contextualize our parameterized model to enhance the module's capacity in exploiting the visiolinguistic associations.Our model outperforms the state-of-the-art NMN model on CLEVR-Ref+ dataset with +8.1% improvement in accuracy on the single-referent test set and +4.3% on the full test set.Additionally, we demonstrate that contextualization provides +11.2% and +1.7% improvements in accuracy over prior NMN models on CLO-SURE and NLVR2.We further evaluate the impact of our contextualization by constructing a contrast set for CLEVR-Ref+, which we call CC-Ref+.We significantly outperform the baselines by as much as +10.4% absolute accuracy on CC-Ref+, illustrating the generalization skills of our approach.Our dataset is publicly available at https://github.com/ McGill-NLP/contextual-nmn.
Arjun R. Akula, Spandana Gella, Keze Wang, Song-Chun Zhu, Siva Reddy
EMNLP (1)2
2020 Words Aren't Enough, Their Order Matters: On the Robustness of Grounding Visual Referring Expressions
abstract
Visual referring expression recognition is a challenging task that requires natural language understanding in the context of an image.We critically examine RefCOCOg, a standard benchmark for this task, using a human study and show that 83.7% of test instances do not require reasoning on linguistic structure, i.e., words are enough to identify the target object, the word order doesn't matter.To measure the true progress of existing models, we split the test set into two sets, one which requires reasoning on linguistic structure and the other which doesn't.Additionally, we create an out-of-distribution dataset Ref-Adv by asking crowdworkers to perturb in-domain examples such that the target object changes.Using these datasets, we empirically show that existing methods fail to exploit linguistic structure and are 12% to 23% lower in performance than the established progress for this task.We also propose two methods, one based on contrastive learning and the other based on multi-task learning, to increase the robustness of ViLBERT, the current state-ofthe-art model for this task.Our datasets are publicly
Arjun R. Akula, Spandana Gella, Yaser Al-Onaizan, Song-Chun Zhu, Siva Reddy
ACL2
2020 An Empirical Study on Robustness to Spurious Correlations using Pre-trained Language Models
abstract
Recent work has shown that pre-trained language models such as BERT improve robustness to spurious correlations in the dataset. Intrigued by these results, we find that the key to their success is generalization from a small amount of counterexamples where the spurious correlations do not hold. When such minority examples are scarce, pre-trained models perform as poorly as models trained from scratch. In the case of extreme minority, we propose to use multi-task learning (MTL) to improve generalization. Our experiments on natural language inference and paraphrase identification show that MTL with the right auxiliary tasks significantly improves performance on challenging examples without hurting the in-distribution performance. Further, we show that the gain from MTL mainly comes from improved generalization from the minority examples. Our results highlight the importance of data diversity for overcoming spurious correlations. 1
Lifu Tu, Garima Lalwani, Spandana Gella, He He 0001
Trans. Assoc. Comput. Linguistics3
2019 Multimodal Abstractive Summarization for How2 Videos
abstract
In this paper, we study abstractive summarization for open-domain videos.Unlike the traditional text news summarization, the goal is less to "compress" text information but rather to provide a fluent textual summary of information that has been collected and fused from different source modalities, in our case video and audio transcripts (or text).We show how a multi-source sequence-to-sequence model with hierarchical attention can integrate information from different modalities into a coherent output, compare various models trained with different modalities and present pilot experiments on the How2 corpus of instructional videos.We also propose a new evaluation metric (Content F1) for abstractive summarization task that measures semantic adequacy rather than fluency of the summaries, which is covered by metrics like ROUGE and BLEU.
Shruti Palaskar, Jindrich Libovický, Spandana Gella, Florian Metze
ACL (1)3
2019 Disambiguating Visual Verbs
abstract
In this article, we introduce a new task, visual sense disambiguation for verbs: given an image and a verb, assign the correct sense of the verb, i.e., the one that describes the action depicted in the image. Just as textual word sense disambiguation is useful for a wide range of NLP tasks, visual sense disambiguation can be useful for multimodal tasks such as image retrieval, image description, and text illustration. We introduce a new dataset, which we call VerSe (short for Verb Sense) that augments existing multimodal datasets (COCO and TUHOI) with verb and sense labels. We explore supervised and unsupervised models for the sense disambiguation task using textual, visual, and multimodal embeddings. We also consider a scenario in which we must detect the verb depicted in an image prior to predicting its sense (i.e., there is no verbal information associated with the image). We find that textual embeddings perform well when gold-standard annotations (object labels and image descriptions) are available, while multimodal embeddings perform well on unannotated images. VerSe is publicly available at https://github.com/spandanagella/verse.
Spandana Gella, Frank Keller, Mirella Lapata
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 A Dataset for Telling the Stories of Social Media Videos
abstract
Video content on social media platforms constitutes a major part of the communication between people, as it allows everyone to share their stories.However, if someone is unable to consume video, either due to a disability or network bandwidth, this severely limits their participation and communication.Automatically telling the stories using multi-sentence descriptions of videos would allow bridging this gap.To learn and evaluate such models, we introduce VideoStory, a new large-scale dataset for video description as a new challenge for multisentence video description.Our VideoStory captions dataset is complementary to prior work and contains 20k videos posted publicly on a social media platform amounting to 396 hours of video with 123k sentences, temporally aligned to the video.* *Work done while SG was intern at Facebook AI Research.
Spandana Gella, Mike Lewis, Marcus Rohrbach
EMNLP1
2017 Image Pivoting for Learning Multilingual Multimodal Representations
abstract
In this paper we propose a model to learn multimodal multilingual representations for matching images and sentences in different languages, with the aim of advancing multilingual versions of image search and image understanding.Our model learns a common representation for images and their descriptions in two different languages (which need not be parallel) by considering the image as a pivot between two languages.We introduce a new pairwise ranking loss function which can handle both symmetric and asymmetric similarity between the two modalities.We evaluate our models on image-description ranking for German and English, and on semantic textual similarity of image descriptions in English.In both cases we achieve state-of-the-art performance.
Spandana Gella, Rico Sennrich, Frank Keller, Mirella Lapata
EMNLP1
2016 Viral Spread via Entertainment and Voice-Messaging Among Telephone Users in India
abstract
We explore how development-related, voice-based, information services could organically spread among low-literate masses in the developing world. We report lessons learned from a remote deployment of "Polly" in India (from the US) to spread job-related information. Polly is an entertainment driven, voice-based service, available over simple phones that is aimed at familiarizing people with speech interfaces and mass-dissemination of development related information to low-literate users. In 2012, Polly had become viral in Pakistan and successfully spread recorded newspaper job ads to thousands of mobile phone users. Remotely deployed in India, Polly did not take off immediately as it did in Pakistan. Instead, it initially entered a six-month long phase of fluctuating, intermittent activity. We experimented with various forms of seeding and it eventually transitioned into a viral phase, with sustained transmission that continued for five months but without (exponential) growth. Finally, interface adjustments in response to user feedback enabling plain-voice asynchronous voice-messaging resulted in an abrupt exponential and viral growth amassing 10,349 phone calls by 1,613 users over a span of seven days. Of these, 299 users also transitioned to the job service. User feedback and surveys suggest possible reasons for each phase. We study the challenges of remote deployment and the interplay of user interface; language of the system; seeding mechanisms and active response to user feedback towards the uptake of the service. We also report a detailed comparison of viral spread in the two countries.
Agha Ali Raza, Rajat Kulshreshtha, Spandana Gella, Sean Olin Blagsvedt, Maya Chandrasekaran, Bhiksha Raj, Ronald Rosenfeld
ICTD3
2016 Unsupervised Visual Sense Disambiguation for Verbs using Multimodal Embeddings
abstract
We introduce a new task, visual sense disambiguation for verbs: given an image and a verb, assign the correct sense of the verb, i.e., the one that describes the action depicted in the image.Just as textual word sense disambiguation is useful for a wide range of NLP tasks, visual sense disambiguation can be useful for multimodal tasks such as image retrieval, image description, and text illustration.We introduce VerSe, a new dataset that augments existing multimodal datasets (COCO and TUHOI) with sense labels.We propose an unsupervised algorithm based on Lesk which performs visual sense disambiguation using textual, visual, or multimodal embeddings.We find that textual embeddings perform well when goldstandard textual annotations (object labels and image descriptions) are available, while multimodal embeddings perform well on unannotated images.We also verify our findings by using the textual and multimodal embeddings as features in a supervised setting and analyse the performance of visual sense disambiguation task.VerSe is made publicly available and can be downloaded at: https://github.com/spandanagella/verse.
Spandana Gella, Mirella Lapata, Frank Keller
HLT-NAACL1
2014 Learning Word Sense Distributions, Detecting Unattested Senses and Identifying Novel Senses Using Topic Models
abstract
Unsupervised word sense disambiguation (WSD) methods are an attractive approach to all-words WSD due to their non-reliance on expensive annotated data.Unsupervised estimates of sense frequency have been shown to be very useful for WSD due to the skewed nature of word sense distributions.This paper presents a fully unsupervised topic modelling-based approach to sense frequency estimation, which is highly portable to different corpora and sense inventories, in being applicable to any part of speech, and not requiring a hierarchical sense inventory, parsing or parallel text.We demonstrate the effectiveness of the method over the tasks of predominant sense learning and sense distribution acquisition, and also the novel tasks of detecting senses which aren't attested in the corpus, and identifying novel senses in the corpus which aren't captured in the sense inventory.
Jey Han Lau, Paul Cook, Diana McCarthy, Spandana Gella, Timothy Baldwin
ACL (1)4
2014 One Sense per Tweeter ... and Other Lexical Semantic Tales of Twitter
abstract
In recent years, microblogs such as Twitter have emerged as a new communication channel.Twitter in particular has become the target of a myriad of content-based applications including trend analysis and event detection, but there has been little fundamental work on the analysis of word usage patterns in this text type.In this paper -inspired by the one-sense-perdiscourse heuristic of Gale et al. ( 1992) -we investigate user-level sense distributions, and detect strong support for "one sense per tweeter".As part of this, we construct a novel sense-tagged lexical sample dataset based on Twitter and a web corpus.
Spandana Gella, Paul Cook, Timothy Baldwin
EACL1
2014 POS Tagging of English-Hindi Code-Mixed Social Media Content
abstract
Code-mixing is frequently observed in user generated content on social media, especially from multilingual users. The linguistic complexity of such content is compounded by presence of spelling vari-ations, transliteration and non-adherance to formal grammar. We describe our initial efforts to create a multi-level an-notated corpus of Hindi-English code-mixed text collated from Facebook fo-rums, and explore language identifica-tion, back-transliteration, normalization and POS tagging of this data. Our re-sults show that language identification and transliteration for Hindi are two major challenges that impact POS tagging accu-racy. 1
Yogarshi Vyas, Spandana Gella, Kalika Bali, Monojit Choudhury
EMNLP2
2014 Mapping WordNet Domains, WordNet Topics and Wikipedia Categories to Generate Multilingual Domain Specific Resources
Spandana Gella, Carlo Strapparava, Vivi Nastase
LREC1