VLDB 2026 Research / reviewers in the wild / expert
Aishwarya Padmakumar
dblp:160/9492
· DBLP profile ↗
14ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AEGIS2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM GuardrailsabstractShaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar, Aishwarya Padmakumar, Traian Rebedea, Jibin Rajan Varghese, Christopher Parisien. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar, Aishwarya Padmakumar, Traian Rebedea, Jibin Rajan Varghese, Christopher Parisien |
NAACL (Long Papers) | 4 |
| 2024 | VLN-Video: Utilizing Driving Videos for Outdoor Vision-and-Language NavigationabstractOutdoor Vision-and-Language Navigation (VLN) requires an agent to navigate through realistic 3D outdoor environments based on natural language instructions. The performance of existing VLN methods is limited by insufficient diversity in navigation environments and limited training data. To address these issues, we propose VLN-Video, which utilizes the diverse outdoor environments present in driving videos in multiple cities in the U.S. augmented with automatically generated navigation instructions and actions to improve outdoor VLN performance. VLN-Video combines the best of intuitive classical approaches and modern deep learning techniques, using template infilling to generate grounded non-repetitive navigation instructions, combined with an image rotation similarity based navigation action predictor to obtain VLN style data from driving videos for pretraining deep learning VLN models. We pre-train the model on the Touchdown dataset and our video-augmented dataset created from driving videos with three proxy tasks: Masked Language Modeling, Instruction and Trajectory Matching, and Next Action Prediction, so as to learn temporally-aware and visually-aligned instruction representations. The learned instruction representation is adapted to the state-of-the-art navigation agent when fine-tuning on the Touchdown dataset. Empirical results demonstrate that VLN-Video significantly outperforms previous state-of-the-art models by 2.1% in task completion rate, achieving a new state-of-the-art on the Touchdown dataset. Jialu Li 0001, Aishwarya Padmakumar, Gaurav S. Sukhatme, Mohit Bansal |
AAAI | 2 |
| 2023 | KILM: Knowledge Injection into Encoder-Decoder Language ModelsabstractYan Xu, Mahdi Namazifar, Devamanyu Hazarika, Aishwarya Padmakumar, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yan Xu 0012, Mahdi Namazifar, Devamanyu Hazarika, Aishwarya Padmakumar, Yang Liu 0004, Dilek Hakkani-Tür |
ACL (1) | 4 |
| 2023 | Multimodal Embodied Plan Prediction Augmented with Synthetic Embodied DialogueabstractEmbodied task completion is a challenge where an agent in a simulated environment must predict environment actions to complete tasks based on natural language instructions and egocentric visual observations.We propose a variant of this problem where the agent predicts actions at a higher level of abstraction called a plan, which helps make agent actions more interpretable and can be obtained from the appropriate prompting of large language models.We show that multimodal transformer models can outperform language-only models for this problem but fall significantly short of oracle plans.Since collecting human-human dialogues for embodied environments is expensive and time-consuming, we propose a method to synthetically generate such dialogues, which we then use as training data for plan prediction.We demonstrate that multimodal transformer models can attain strong zero-shot performance from our synthetic data, outperforming language-only models trained on humanhuman data. Aishwarya Padmakumar, Mert Inan, Spandana Gella, Patrick Lange, Dilek Hakkani-Tür |
EMNLP | 1 |
| 2022 | TEACh: Task-Driven Embodied Agents That ChatabstractRobots operating in human spaces must be able to engage in natural language interaction, both understanding and executing instructions, and using conversation to resolve ambiguity and correct mistakes. To study this, we introduce TEACh, a dataset of over 3,000 human-human, interactive dialogues to complete household tasks in simulation. A Commander with access to oracle information about a task communicates in natural language with a Follower. The Follower navigates through and interacts with the environment to complete tasks varying in complexity from "Make Coffee" to "Prepare Breakfast", asking questions and getting additional information from the Commander. We propose three benchmarks using TEACh to study embodied intelligence challenges, and we evaluate initial models' abilities in dialogue understanding, language grounding, and task execution. Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gökhan Tür, Dilek Hakkani-Tür |
AAAI | 1 |
| 2022 | ALFRED-L: Investigating the Role of Language for Action Learning in Interactive Visual EnvironmentsabstractArjun Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tur. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Arjun R. Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tür |
EMNLP | 3 |
| 2022 | Dialog Acts for Task Driven Embodied AgentsabstractEmbodied agents need to be able to interact in natural language -understanding task descriptions and asking appropriate follow up questions to obtain necessary information to be effective at successfully accomplishing tasks for a wide range of users.In this work, we propose a set of dialog acts for modelling such dialogs and annotate the TEACh dataset that includes over 3,000 situated, task oriented conversations (consisting of 39.5k utterances in total) with dialog acts.TEACh-DA is one of the first large scale dataset of dialog act annotations for embodied task completion.Furthermore, we demonstrate the use of this annotated dataset in training models for tagging the dialog acts of a given utterance, predicting the dialog act of the next response given a dialog history, and use the dialog acts to guide agent's non-dialog behaviour.In particular, our experiments on the TEACh Execution from Dialog History task where the model predicts the sequence of low level actions to be executed in the environment for embodied task completion, demonstrate that dialog acts can improve end task success rate by up to 2 points compared to the system without dialog acts. Spandana Gella, Aishwarya Padmakumar, Patrick Lange, Dilek Hakkani-Tür |
SIGDIAL | 2 |
| 2021 | Dialog Policy Learning for Joint Clarification and Active Learning QueriesabstractIntelligent systems need to be able to recover from mistakes, resolve uncertainty, and adapt to novel concepts not seen during training. Dialog interaction can enable this by the use of clarifications for correction and resolving uncertainty, and active learning queries to learn new concepts encountered during operation. Prior work on dialog systems has either focused on exclusively learning how to perform clarification/ information seeking, or to perform active learning. In this work, we train a hierarchical dialog policy to jointly perform {\it both} clarification and active learning in the context of an interactive language-based image retrieval task motivated by an online shopping application, and demonstrate that jointly learning dialog policies for clarification and active learning is more effective than the use of static dialog policies for one or both of these functions. Aishwarya Padmakumar, Raymond J. Mooney |
AAAI | 1 |
| 2021 | Generative Conversational NetworksabstractAlexandros Papangelis, Karthik Gopalakrishnan, Aishwarya Padmakumar, Seokhwan Kim, Gokhan Tur, Dilek Hakkani-Tur. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021. Alexandros Papangelis, Karthik Gopalakrishnan 0001, Aishwarya Padmakumar, Seokhwan Kim, Gökhan Tür, Dilek Hakkani-Tür |
SIGDIAL | 3 |
| 2020 | Jointly Improving Parsing and Perception for Natural Language Commands through Human-Robot DialogabstractIn this work, we present methods for using human-robot dialog to improve language understanding for a mobile robot agent. The agent parses natural language to underlying semantic meanings and uses robotic sensors to create multi-modal models of perceptual concepts like red and heavy. The agent can be used for showing navigation routes, delivering objects to people, and relocating objects from one location to another. We use dialog clari_cation questions both to understand commands and to generate additional parsing training data. The agent employs opportunistic active learning to select questions about how words relate to objects, improving its understanding of perceptual concepts. We evaluated this agent on Amazon Mechanical Turk. After training on data induced from conversations, the agent reduced the number of dialog questions it asked while receiving higher usability ratings. Additionally, we demonstrated the agent on a robotic platform, where it learned new perceptual concepts on the y while completing a real-world task. Jesse Thomason, Aishwarya Padmakumar, Jivko Sinapov, Nick Walker 0001, Yuqian Jiang, Harel Yedidsion, Justin W. Hart, Peter Stone 0001, Raymond J. Mooney |
J. Artif. Intell. Res. | 2 |
| 2019 | Improving Grounded Natural Language Understanding through Human-Robot DialogabstractNatural language understanding for robotics can require substantial domain- and platform-specific engineering. For example, for mobile robots to pick-and-place objects in an environment to satisfy human commands, we can specify the language humans use to issue such commands, and connect concept words like red can to physical object properties. One way to alleviate this engineering for a new domain is to enable robots in human environments to adapt dynamically-continually learning new language constructions and perceptual concepts. In this work, we present an end-to-end pipeline for translating natural language commands to discrete robot actions, and use clarification dialogs to jointly improve language parsing and concept grounding. We train and evaluate this agent in a virtual setting on Amazon Mechanical Turk, and we transfer the learned agent to a physical robot platform to demonstrate it in the real world. Jesse Thomason, Aishwarya Padmakumar, Jivko Sinapov, Nick Walker 0001, Yuqian Jiang, Harel Yedidsion, Justin W. Hart, Peter Stone 0001, Raymond J. Mooney |
ICRA | 2 |
| 2018 | Learning a Policy for Opportunistic Active LearningabstractActive learning identifies data points to label that are expected to be the most useful in improving a supervised model.Opportunistic active learning incorporates active learning into interactive tasks that constrain possible queries during interactions.Prior work has shown that opportunistic active learning can be used to improve grounding of natural language descriptions in an interactive object retrieval task.In this work, we use reinforcement learning for such an object retrieval task, to learn a policy that effectively trades off task completion with model improvement that would benefit future tasks. Aishwarya Padmakumar, Peter Stone 0001, Raymond J. Mooney |
EMNLP | 1 |
| 2017 | Integrated Learning of Dialog Strategies and Semantic ParsingabstractNatural language understanding and dialog management are two integral components of interactive dialog systems.Previous research has used machine learning techniques to individually optimize these components, with different forms of direct and indirect supervision.We present an approach to integrate the learning of both a dialog strategy using reinforcement learning, and a semantic parser for robust natural language understanding, using only natural dialog interaction for supervision.Experimental results on a simulated task of robot instruction demonstrate that joint learning of both components improves dialog performance over learning either of these components alone. Aishwarya Padmakumar, Jesse Thomason, Raymond J. Mooney |
EACL (1) | 1 |
| 2015 | Automated Linguistic Personalization of Targeted Marketing Messages Mining User-Generated Text on Social Media
Rishiraj Saha Roy, Aishwarya Padmakumar, Guna Prasaad Jeganathan, Ponnurangam Kumaraguru |
CICLing (2) | 2 |