EDBT 2026 Demo / reviewers in the wild / expert
Jesse Thomason
dblp:130/2863
· DBLP profile ↗
42ranked-venue papers
7as first author
27since 2021 · last 2026
0000-0001-9199-0633ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 7 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 4 since 2021Systems, architecture and hardware · 7 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning to Deliberate: Meta-policy Collaboration for Agentic LLMs with Multi-agent Reinforcement LearningabstractMulti-agent systems of large language models (LLMs) show promise for complex reasoning, but their effectiveness is often limited by fixed collaboration protocols. These frameworks typically focus on macro-level orchestration while overlooking agents’ internal deliberative capabilities. This critical meta-cognitive blindspot treats agents as passive executors unable to adapt their strategy based on internal cognitive states like uncertainty or confidence. We introduce the Meta-Policy Deliberation Framework (MPDF), where agents learn a decentralized policy over a set of high-level meta-cognitive actions: Persist, Refine, and Concede. To overcome the instability of traditional policy gradients in this setting, we develop SoftRankPO, a novel reinforcement learning algorithm. SoftRankPO stabilizes training by shaping advantages based on the rank of rewards mapped through smooth normal quantiles, making the learning process robust to reward variance. Experiments show that MPDF with SoftRankPO achieves a 4-5% absolute gain in average accuracy across six mathematical and general reasoning benchmarks compared to state-of-the-art heuristic and learning-based multi-agent reasoning algorithms. Our work presents a paradigm for learning adaptive, meta-cognitive policies for multi-agent LLM systems, shifting the focus from designing fixed protocols to learning dynamic, deliberative strategies. Jesse Thomason |
AAAI | 2 |
| 2026 | Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model ExplanationsabstractKeyu He, Tejas Srinivasan, Brihi Joshi, Xiang Ren, Jesse Thomason, Swabha Swayamdipta. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Keyu He, Tejas Srinivasan, Brihi Joshi, Xiang Ren 0001, Jesse Thomason, Swabha Swayamdipta |
ACL (1) | 5 |
| 2026 | Adjust for Trust: Mitigating Trust-Induced Inappropriate Reliance on AI AssistanceabstractUser trust biases reliance on AI assistance during decision-making tasks, with overly low and high levels of trust resulting in increased under- and over-reliance, respectively. We propose that AI assistants should adapt their behavior through trust-adaptive interventions to mitigate such inappropriate reliance. For instance, when user trust is low, that providing an AI explanation can elicit more careful consideration of the assistant’s advice by the user. In two decision-making scenarios—laypeople answering science questions and medical doctors making diagnoses—we find that providing supporting and counter-explanations during moments of low and high trust, respectively, yields up to 38% reduction in inappropriate reliance and 20% improvement in decision accuracy. We are similarly able to reduce over-reliance by adaptively inserting forced pauses to promote deliberation. Our results highlight how AI adaptation to user trust facilitates appropriate reliance, presenting exciting avenues for improving human-AI collaboration. Tejas Srinivasan, Jesse Thomason |
IUI | 2 |
| 2026 | DiPS: Dialogue Policy Selection for High-Stakes Persuasion AgentsabstractLarge Language Models (LLMs) often struggle with persuasion in high-stakes scenarios. People’s individual personalities and concerns require tailored strategies rather than a one-size-fits-all approach. To address this challenge, we focus on a fire-rescue scenario in which an operator must persuade a resident to evacuate as a high-stakes persuasion domain and propose Dialogue Policy Selection (DiPS), a Q-learning framework to dynamically select persuasion strategies adapted to the evolving conversational context. Specifically, we train a critic, trained to maximize the chance of evacuation success, to select a persuasion policy at each turn based on the resident’s recent utterances. We then evaluate DiPS against multiple baselines in both simulated and real human interactions. We find that DiPS achieves higher evacuation success than a zero-shot LLM and generic RAG-augmented approach. Mousumi Das, Abrar Anwar, Jesse Thomason, David Traum |
SIGDIAL | 4 |
| 2025 | Promoting Cognitive Health in Elder Care with Large Language Model-Powered Socially Assistive RobotsabstractAs the global population ages, there is increasing need for accessible technologies that promote cognitive health and detect early signs of cognitive decline. This research demonstrates the potential for in-residence monitoring and assessment of cognitive health using large language model (LLM)-powered socially assistive robots (SARs). We conducted a 5-week within-subjects study involving 22 older adults in retirement homes to investigate the feasibility of large language model (LLM)-powered socially assistive robots (SARs) for promoting and assessing cognitive health. We designed tasks that involved verbal dialogue based on clinically validated cognitive tools. Our findings reveal improved task performance after three robot-administered sessions, with significantly more detailed picture descriptions, fewer word repetitions in semantic fluency, and reduced need for hints. We found that older adults were more socially engaged in robot-administered tasks compared to those administered by a human, and they accepted and were willing to engage with socially assistive robots (SARs) in this context, which had not been tested before. Maria R. Lima, Amy O'Connell, Feiyang Zhou, Alethea Nagahara, Avni Hulyalkar, Anura Deshpande, Jesse Thomason, Ravi Vaidyanathan, Maja J. Mataric |
CHI | 7 |
| 2025 | Why Do Some Inputs Break Low-Bit LLM Quantization?abstractLow-bit weight-only quantization significantly reduces the memory footprint of large language models (LLMs), but disproportionately affects certain examples.We analyze diverse 3-4 bit methods on LLMs ranging from 7B-70B in size and find that the quantization errors of 50 pairs of methods are strongly correlated (avg.ρ = 0.82) on FineWeb examples.Moreover, the residual stream magnitudes of full-precision models are indicative of future quantization errors.We further establish a hypothesis that relates the residual stream magnitudes to error amplification and accumulation over layers.Using LLM localization techniques, early exiting, and activation patching, we show that examples with large errors rely on precise residual activations in the later layers, and that the outputs of MLP gates play a crucial role in maintaining the perplexity.Our work reveals why certain examples result in large quantization errors and which model components are most critical for performance preservation. Ting-Yun Chang, Muru Zhang, Jesse Thomason, Robin Jia |
EMNLP | 3 |
| 2025 | Large Language Models Do Multi-Label Classification DifferentlyabstractMulti-label classification is prevalent in realworld settings, but the behavior of Large Language Models (LLMs) in this setting is understudied.We investigate how autoregressive LLMs perform multi-label classification, focusing on subjective tasks, by analyzing the output distributions of the models at each label generation step.We find that the initial probability distribution for the first label often does not reflect the eventual final output, even in terms of relative order and find LLMs tend to suppress all but one label at each generation step.We further observe that as model scale increases, their token distributions exhibit lower entropy and higher single-label confidence, but the internal relative ranking of the labels improves.Finetuning methods such as supervised finetuning and reinforcement learning amplify this phenomenon.We introduce the task of distribution alignment for multi-label settings: aligning LLM-derived label distributions with empirical distributions estimated from annotator responses in subjective tasks.We propose both zero-shot and supervised methods which improve both alignment and predictive performance over existing approaches.We find one method -taking the max probability over all label generation distributions instead of just using the initial probability distribution -improves both distribution alignment and overall F1 classification without adding any additional computation. Marcus Ma, Georgios Chochlakis, Niyantha Maruthu Pandiyan, Jesse Thomason, Shri Narayanan |
EMNLP | 4 |
| 2025 | Multimodal Synthetic Data Finetuning and Model Collapse: Insights from VLMs and Diffusion Models
Zizhao Hu, Jesse Thomason |
ICMI | 3 |
| 2025 | Language Models Can Infer Action Semantics for Symbolic Planners from Environment FeedbackabstractWang Bill Zhu, Ishika Singh, Robin Jia, Jesse Thomason. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Wang Zhu 0001, Ishika Singh, Robin Jia, Jesse Thomason |
NAACL (Long Papers) | 4 |
| 2024 | When Parts Are Greater Than Sums: Individual LLM Components Can Outperform Full ModelsabstractThis paper studies in-context learning by decomposing the output of large language models into the individual contributions of attention heads and MLPs (components).We observe curious components: good-performing ones that individually do well on a classification task, even when the full model performs poorly; bad-performing ones that do much worse than chance; and label-biased components that always predict the same label.We find that component accuracies are well-correlated across different demonstration sets and perturbations of prompt templates.Based on our findings, we propose component reweighting, which learns to linearly re-scale the component activations from a few labeled examples.Given 24 labeled examples, our method improves by an average of 6.0% accuracy points over 24-shot ICL across 8 tasks on Llama-2-7B.Overall, this paper both enriches our understanding of ICL and provides a practical method for improvement by examining model internals.Template 1 {text}\nIs this a piece of news regarding World, Sports, Business, or Technology?Template 2 {text} Is this a piece of news regarding World, Sports, Business, or Technology?Correlation = 0.81 Ting-Yun Chang, Jesse Thomason, Robin Jia |
EMNLP | 2 |
| 2024 | ViSaRL: Visual Reinforcement Learning Guided by Human SaliencyabstractTraining robots to perform complex control tasks from high-dimensional pixel input using reinforcement learning (RL) is sample-inefficient, because image observations are comprised primarily of task-irrelevant information. By contrast, humans are able to visually attend to task-relevant objects and areas. Based on this insight, we introduce Visual Saliency-Guided Reinforcement Learning (ViSaRL). Using ViSaRL to learn visual representations significantly improves the success rate, sample efficiency, and generalization of an RL agent on diverse tasks including DeepMind Control benchmark, robot manipulation in simulation and on a real robot. We present approaches for incorporating saliency into both CNN and Transformer-based encoders. We show that visual representations learned using ViSaRL are robust to various sources of visual perturbations including perceptual noise and scene variations. ViSaRL nearly doubles success rate on the real-robot tasks compared to the baseline which does not use saliency. Anthony Liang, Jesse Thomason, Erdem Biyik |
IROS | 2 |
| 2024 | Efficient End-to-End Visual Document Understanding with Rationale DistillationabstractWang Zhu, Alekh Agarwal, Mandar Joshi, Robin Jia, Jesse Thomason, Kristina Toutanova. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Wang Zhu 0001, Alekh Agarwal, Mandar Joshi, Robin Jia, Jesse Thomason, Kristina Toutanova |
NAACL-HLT | 5 |
| 2024 | Do Localization Methods Actually Localize Memorized Data in LLMs? A Tale of Two BenchmarksabstractTing-Yun Chang, Jesse Thomason, Robin Jia. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Ting-Yun Chang, Jesse Thomason, Robin Jia |
NAACL-HLT | 2 |
| 2024 | Which One? Leveraging Context Between Objects and Multiple Views for Language GroundingabstractChancharik Mitra, Abrar Anwar, Rodolfo Corona, Dan Klein, Trevor Darrell, Jesse Thomason. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chancharik Mitra, Abrar Anwar, Rodolfo Corona, Daniel Klein 0001, Trevor Darrell, Jesse Thomason |
NAACL-HLT | 6 |
| 2023 | The Sem-Lex Benchmark: Modeling ASL Signs and their PhonemesabstractSign language recognition and translation technologies have the potential to increase access and inclusion of deaf signing communities, but research progress is bottlenecked by a lack of representative data. We introduce a new resource for American Sign Language (ASL) modeling, the Sem-Lex Benchmark. The Benchmark is the current largest of its kind, consisting of over 84k videos of isolated sign productions from deaf ASL signers who gave informed consent and received compensation. Human experts aligned these videos with other sign language resources including ASL-LEX, SignBank, and ASL Citizen, enabling useful expansions for sign and phonological feature recognition. We present a suite of experiments which make use of the linguistic information in ASL-LEX, evaluating the practicality and fairness of the Sem-Lex Benchmark for isolated sign recognition (ISR). We use an SL-GCN model to show that the phonological features are recognizable with 85% accuracy, and that they are effective as an auxiliary target to ISR. Learning to recognize phonological features alongside gloss results in a 6% improvement for few-shot ISR accuracy and a 2% improvement for ISR accuracy overall. Instructions for downloading the data can be found at https://github.com/leekezar/SemLex. Lee Kezar, Jesse Thomason, Naomi Caselli, Zed Sevcikova Sehyr, Elana Pontecorvo |
ASSETS | 2 |
| 2023 | Iterative Vision-and-Language NavigationabstractWe present Iterative Vision-and-Language Navigation (IVLN), a paradigm for evaluating language-guided agents navigating in a persistent environment over time. Existing Vision-and-Language Navigation (VLN) benchmarks erase the agent's memory at the beginning of every episode, testing the ability to perform cold-start navigation with no prior information. However, deployed robots occupy the same environment for long periods of time. The IVLN paradigm addresses this disparity by training and evaluating VLN agents that maintain memory across tours of scenes that consist of up to 100 ordered instruction-following Room-to-Room (R2R) episodes, each defined by an individual language instruction and a target path. We present discrete and continuous Iterative Room-to-Room (IR2R) benchmarks comprising about 400 tours each in 80 indoor scenes. We find that extending the implicit memory of high-performing transformer VLN agents is not sufficient for IVLN, but agents that build maps can benefit from environment persistence, motivating a renewed focus on map-building agents in VLN. Jacob Krantz, Shurjo Banerjee, Wang Zhu 0001, Jason J. Corso, Stefan Lee, Jesse Thomason |
CVPR | 7 |
| 2023 | Improving Sign Recognition with PhonologyabstractWe use insights from research on American Sign Language (ASL) phonology to train models for isolated sign language recognition (ISLR), a step towards automatic sign language understanding.Our key insight is to explicitly recognize the role of phonology in sign production to achieve more accurate ISLR than existing work which does not consider sign language phonology.We train ISLR models that take in pose estimations of a signer producing a single sign to predict not only the sign but additionally its phonological characteristics, such as the handshape.These auxiliary predictions lead to a nearly 9% absolute gain in sign recognition accuracy on the WLASL benchmark, with consistent improvements in ISLR regardless of the underlying prediction model architecture.This work has the potential to accelerate linguistic research in the domain of signed languages and reduce communication barriers between deaf and hearing people. Lee Kezar, Jesse Thomason, Zed Sevcikova Sehyr |
EACL | 2 |
| 2023 | Chain-of-Questions Training with Latent Answers for Robust Multistep Question AnsweringabstractWe propose Chain-of-Questions, a framework that trains a model to robustly answer multistep questions by generating and answering sub-questions.We obtain supervision for subquestions from human-annotated question decomposition meaning representation (QDMR), but QDMR does not include annotated answers to sub-questions.To overcome this technical challenge, we treat sub-answers as latent variables and infer them with a novel dynamic mixture of Hard-EM and MAPO.Chain-of-Questions is effective and robust, greatly outperforming strong neuro-symbolic methods by 9.0 F1 on a DROP contrast set and GPT-3.5 by 24.3 F1 on a HOTPOTQA adversarial set. Wang Zhu 0001, Jesse Thomason, Robin Jia |
EMNLP | 2 |
| 2023 | Exploring Strategies for Modeling Sign Language PhonologyabstractLike speech, signs are composed of discrete, recombinable features called phonemes.Prior work shows that models which can recognize phonemes are better at sign recognition, motivating deeper exploration into strategies for modeling sign language phonemes.In this work, we learn graph convolution networks to recognize the sixteen phoneme "types" found in ASL-LEX 2.0.Specifically, we explore how learning strategies like multi-task and curriculum learning can leverage mutually useful information between phoneme types to facilitate better modeling of sign language phonemes.Results on the Sem-Lex Benchmark show that curriculum learning yields an average accuracy of 87% across all phoneme types, outperforming fine-tuning and multi-task strategies for most phoneme types. Lee Kezar, Tejas Srinivasan, Riley Carlin, Jesse Thomason, Zed Sevcikova Sehyr, Naomi Caselli |
ESANN | 4 |
| 2023 | ProgPrompt: Generating Situated Robot Task Plans using Large Language ModelsabstractTask planning can require defining myriad domain knowledge about the world in which a robot needs to act. To ameliorate that effort, large language models (LLMs) can be used to score potential next actions during task planning, and even generate action sequences directly, given an instruction in natural language with no additional domain information. However, such methods either require enumerating all possible next steps for scoring, or generate free-form text that may contain actions not possible on a given robot in its current context. We present a programmatic LLM prompt structure that enables plan generation functional across situated environments, robot capabilities, and tasks. Our key insight is to prompt the LLM with program-like specifications of the available actions and objects in an environment, as well as with example programs that can be executed. We make concrete recommendations about prompt structure and generation constraints through ablation experiments, demonstrate state of the art success rates in VirtualHome household tasks, and deploy our method on a physical robot arm for tabletop tasks. Website at progprompt.github.io Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal 0001, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, Animesh Garg |
ICRA | 8 |
| 2023 | Multimodal Speech Recognition for Language-Guided Embodied AgentsabstractBenchmarks for language-guided embodied agents typically assume text-based instructions, but deployed agents will encounter spoken instructions.While Automatic Speech Recognition (ASR) models can bridge the input gap, erroneous ASR transcripts can hurt the agents' ability to complete tasks.We propose training a multimodal ASR model that utilizes the accompanying visual context to reduce errors in spoken instruction transcripts.We train our model on a dataset of synthetic spoken instructions, derived from the ALFRED household task dataset, where we simulate acoustic noise by systematically masking spoken words.We find that utilizing visual observations facilitates masked word recovery, with multimodal ASR models recovering up to 30% more masked words than unimodal baselines.We also find that spoken instructions transcribed by multimodal ASR models result in higher task completion success rates for a language-guided embodied agent.github.com Allen Chang, Xiaoyuan Zhu, Aarav Monga, Seoho Ahn, Tejas Srinivasan, Jesse Thomason |
INTERSPEECH | 6 |
| 2023 | RREx-BoT: Remote Referring Expressions with a Bag of TricksabstractHousehold robots operate in the same space for years. Such robots incrementally build dynamic maps that can be used for tasks requiring remote object localization. However, benchmarks in robot learning often test generalization through inference on tasks in unobserved environments. In an observed environment, locating an object is reduced to choosing from among all object proposals in the environment, which may number in the 100,000s. Armed with this intuition, using only a generic vision-language scoring model with minor modifications for 3d encoding and operating in an embodied environment, we demonstrate an absolute performance gain of 9.84% on remote object grounding above state of the art models for REVERIE and of 5.04% on FAO. When allowed to pre-explore an environment, we also exceed the previous state of the art pre-exploration method on REVERIE. Additionally, we demonstrate our model on a real-world TurtleBot platform, highlighting the simplicity and usefulness of the approach. Our analysis outlines a “bag of tricks” essential for accomplishing this task, from utilizing 3d coordinates and context, to gener-alizing vision-language models to large 3d search spaces. Gunnar A. Sigurdsson, Jesse Thomason, Gaurav S. Sukhatme, Robinson Piramuthu |
IROS | 2 |
| 2022 | TEACh: Task-Driven Embodied Agents That ChatabstractRobots operating in human spaces must be able to engage in natural language interaction, both understanding and executing instructions, and using conversation to resolve ambiguity and correct mistakes. To study this, we introduce TEACh, a dataset of over 3,000 human-human, interactive dialogues to complete household tasks in simulation. A Commander with access to oracle information about a task communicates in natural language with a Follower. The Follower navigates through and interacts with the environment to complete tasks varying in complexity from "Make Coffee" to "Prepare Breakfast", asking questions and getting additional information from the Commander. We propose three benchmarks using TEACh to study embodied intelligence challenges, and we evaluate initial models' abilities in dialogue understanding, language grounding, and task execution. Aishwarya Padmakumar, Jesse Thomason, Ayush Shrivastava, Patrick Lange, Anjali Narayan-Chen, Spandana Gella, Robinson Piramuthu, Gökhan Tür, Dilek Hakkani-Tür |
AAAI | 2 |
| 2022 | Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future DirectionsabstractA long-term goal of AI research is to build intelligent agents that can communicate with humans in natural language, perceive the environment, and perform real-world tasks.Visionand-Language Navigation (VLN) is a fundamental and interdisciplinary research topic towards this goal, and receives increasing attention from natural language processing, computer vision, robotics, and machine learning communities.In this paper, we review contemporary studies in the emerging field of VLN, covering tasks, evaluation metrics, methods, etc.Through structured analysis of current progress and challenges, we highlight the limitations of current VLN and opportunities for future work.This paper serves as a thorough reference for the VLN research community.1 Eliana Stefani, Qi Wu 0001, Jesse Thomason, Xin Wang 0061 |
ACL (1) | 4 |
| 2022 | ALFRED-L: Investigating the Role of Language for Action Learning in Interactive Visual EnvironmentsabstractArjun Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tur. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Arjun R. Akula, Spandana Gella, Aishwarya Padmakumar, Mahdi Namazifar, Mohit Bansal, Jesse Thomason, Dilek Hakkani-Tür |
EMNLP | 6 |
| 2022 | SSL Enables Learning from Sparse Rewards in Image-Goal Navigation
Arjun Majumdar, Gunnar A. Sigurdsson, Robinson Piramuthu, Jesse Thomason, Dhruv Batra, Gaurav S. Sukhatme |
ICML | 4 |
| 2022 | CLiMB: A Continual Learning Benchmark for Vision-and-Language TasksabstractCurrent state-of-the-art vision-and-language models are evaluated on tasks either individually or in a multi-task setting, overlooking the challenges of continually learning (CL) tasks as they arrive. Existing CL benchmarks have facilitated research on task adaptation and mitigating "catastrophic forgetting", but are limited to vision-only and language-only tasks. We present CLiMB, a benchmark to study the challenge of learning multimodal tasks in a CL setting, and to systematically evaluate how upstream continual learning can rapidly generalize to new multimodal and unimodal tasks. CLiMB includes implementations of several CL algorithms and a modified Vision-Language Transformer (ViLT) model that can be deployed on both multimodal and unimodal tasks. We find that common CL methods can help mitigate forgetting during multimodal task learning, but do not enable cross-task knowledge transfer. We envision that CLiMB will facilitate research on a new class of CL algorithms for this challenging multimodal setting. Tejas Srinivasan, Ting-Yun Chang, Leticia Leonor Pinto Alva, Georgios Chochlakis, Jesse Thomason |
NeurIPS | 6 |
| 2020 | ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday TasksabstractWe present ALFRED (Action Learning From Realistic Environments and Directives), a benchmark for learning a mapping from natural language instructions and egocentric vision to sequences of actions for household tasks. ALFRED includes long, compositional tasks with non-reversible state changes to shrink the gap between research benchmarks and real-world applications. ALFRED consists of expert demonstrations in interactive visual environments for 25k natural language directives. These directives contain both high-level goals like “Rinse off a mug and place it in the coffee maker.” and low-level language instructions like “Walk to the coffee maker on the right.” ALFRED tasks are more complex in terms of sequence length, action space, and language than existing vision- and-language task datasets. We show that a baseline model based on recent embodied vision-and-language tasks performs poorly on ALFRED, suggesting that there is significant room for developing innovative grounded visual language understanding models with this benchmark. Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, Dieter Fox |
CVPR | 2 |
| 2020 | Experience Grounds LanguageabstractYonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, Joseph Turian. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Y. Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, Joseph P. Turian |
EMNLP (1) | 3 |
| 2020 | Jointly Improving Parsing and Perception for Natural Language Commands through Human-Robot DialogabstractIn this work, we present methods for using human-robot dialog to improve language understanding for a mobile robot agent. The agent parses natural language to underlying semantic meanings and uses robotic sensors to create multi-modal models of perceptual concepts like red and heavy. The agent can be used for showing navigation routes, delivering objects to people, and relocating objects from one location to another. We use dialog clari_cation questions both to understand commands and to generate additional parsing training data. The agent employs opportunistic active learning to select questions about how words relate to objects, improving its understanding of perceptual concepts. We evaluated this agent on Amazon Mechanical Turk. After training on data induced from conversations, the agent reduced the number of dialog questions it asked while receiving higher usability ratings. Additionally, we demonstrated the agent on a robotic platform, where it learned new perceptual concepts on the y while completing a real-world task. Jesse Thomason, Aishwarya Padmakumar, Jivko Sinapov, Nick Walker 0001, Yuqian Jiang, Harel Yedidsion, Justin W. Hart, Peter Stone 0001, Raymond J. Mooney |
J. Artif. Intell. Res. | 1 |
| 2019 | Prospection: Interpretable plans from language by predicting the futureabstractHigh-level human instructions often correspond to behaviors with multiple implicit steps. In order for robots to be useful in the real world, they must be able to to reason over both motions and intermediate goals implied by human instructions. In this work, we propose a framework for learning representations that convert from a natural-language command to a sequence of intermediate goals for execution on a robot. A key feature of this framework is prospection, training an agent not just to correctly execute the prescribed command, but to predict a horizon of consequences of an action before taking it. We demonstrate the fidelity of plans generated by our framework when interpreting real, crowd-sourced natural language commands for a robot in simulated scenes. Chris Paxton 0001, Yonatan Bisk, Jesse Thomason, Arunkumar Byravan, Dieter Fox |
ICRA | 3 |
| 2019 | Improving Grounded Natural Language Understanding through Human-Robot DialogabstractNatural language understanding for robotics can require substantial domain- and platform-specific engineering. For example, for mobile robots to pick-and-place objects in an environment to satisfy human commands, we can specify the language humans use to issue such commands, and connect concept words like red can to physical object properties. One way to alleviate this engineering for a new domain is to enable robots in human environments to adapt dynamically-continually learning new language constructions and perceptual concepts. In this work, we present an end-to-end pipeline for translating natural language commands to discrete robot actions, and use clarification dialogs to jointly improve language parsing and concept grounding. We train and evaluate this agent in a virtual setting on Amazon Mechanical Turk, and we transfer the learned agent to a physical robot platform to demonstrate it in the real world. Jesse Thomason, Aishwarya Padmakumar, Jivko Sinapov, Nick Walker 0001, Yuqian Jiang, Harel Yedidsion, Justin W. Hart, Peter Stone 0001, Raymond J. Mooney |
ICRA | 1 |
| 2019 | Augmenting Knowledge through Statistical, Goal-oriented Human-Robot DialogabstractSome robots can interact with humans using natural language, and identify service requests through human-robot dialog. However, few robots are able to improve their language capabilities from this experience. In this paper, we develop a dialog agent for robots that is able to interpret user commands using a semantic parser, while asking clarification questions using a probabilistic dialog manager. This dialog agent is able to augment its knowledge base and improve its language capabilities by learning from dialog experiences, e.g., adding new entities and learning new ways of referring to existing entities. We have extensively evaluated our dialog system in simulation as well as with human participants through MTurk and real-robot platforms. We demonstrate that our dialog agent performs better in efficiency and accuracy in comparison to baseline learning agents. Demo video can be found at https://youtu.be/DFB3jbHBqYE. Saeid Amiri, Sujay Bajracharya, Cihangir Goktolga, Jesse Thomason, Shiqi Zhang 0001 |
IROS | 4 |
| 2019 | Improving Robot Success Detection using Static Object DataabstractWe use static object data to improve success detection for stacking objects on and nesting objects in one another. Such actions are necessary for certain robotics tasks, e.g., clearing a dining table or packing a warehouse bin. However, using an RGB-D camera to detect success can be insufficient: same-colored objects can be difficult to differentiate, and reflective silverware cause noisy depth camera perception. We show that adding static data about the objects themselves improves the performance of an end-to-end pipeline for classifying action outcomes. Images of the objects, and language expressions describing them, encode prior geometry, shape, and size information that refine classification accuracy. We collect over 13 hours of egocentric manipulation data for training a model to reason about whether a robot successfully placed unseen objects in or on one another. The model achieves up to a 57% absolute gain over the task baseline on pairs of previously unseen objects. Rosario Scalise, Jesse Thomason, Yonatan Bisk, Siddhartha S. Srinivasa |
IROS | 2 |
| 2018 | Maximum-Variance Total Variation Denoising for Interpretable Spatial Smoothing
Wesley Tansey, Jesse Thomason, James G. Scott |
AAAI | 2 |
| 2018 | Guiding Exploratory Behaviors for Multi-Modal Grounding of Linguistic DescriptionsabstractA major goal of grounded language learning research is to enable robots to connect language predicates to a robot's physical interactive perception of the world. Coupling object exploratory behaviors such as grasping, lifting, and looking with multiple sensory modalities (e.g., audio, haptics, and vision) enables a robot to ground non-visual words like ``heavy'' as well as visual words like ``red''. A major limitation of existing approaches to multi-modal language grounding is that a robot has to exhaustively explore training objects with a variety of actions when learning a new such language predicate. This paper proposes a method for guiding a robot's behavioral exploration policy when learning a novel predicate based on known grounded predicates and the novel predicate's linguistic relationship to them. We demonstrate our approach on two datasets in which a robot explored large sets of objects and was tasked with learning to recognize whether novel words applied to those objects. Jesse Thomason, Jivko Sinapov, Raymond J. Mooney, Peter Stone 0001 |
AAAI | 1 |
| 2018 | Multi-modal Predicate Identification using Dynamically Learned Robot ControllersabstractIntelligent robots frequently need to explore the objects in their working environments. Modern sensors have enabled robots to learn object properties via perception of multiple modalities. However, object exploration in the real world poses a challenging trade-off between information gains and exploration action costs. Mixed observability Markov decision process (MOMDP) is a framework for planning under uncertainty, while accounting for both fully and partially observable components of the state. Robot perception frequently has to face such mixed observability. This work enables a robot equipped with an arm to dynamically construct query-oriented MOMDPs for multi-modal predicate identification (MPI) of objects. The robot's behavioral policy is learned from two datasets collected using real robots. Our approach enables a robot to explore object properties in a way that is significantly faster while improving accuracies in comparison to existing methods that rely on hand-coded exploration strategies. Saeid Amiri, Suhua Wei, Shiqi Zhang 0001, Jivko Sinapov, Jesse Thomason, Peter Stone 0001 |
IJCAI | 5 |
| 2017 | Integrated Learning of Dialog Strategies and Semantic ParsingabstractNatural language understanding and dialog management are two integral components of interactive dialog systems.Previous research has used machine learning techniques to individually optimize these components, with different forms of direct and indirect supervision.We present an approach to integrate the learning of both a dialog strategy using reinforcement learning, and a semantic parser for robust natural language understanding, using only natural dialog interaction for supervision.Experimental results on a simulated task of robot instruction demonstrate that joint learning of both components improves dialog performance over learning either of these components alone. Aishwarya Padmakumar, Jesse Thomason, Raymond J. Mooney |
EACL (1) | 2 |
| 2017 | Multi-Modal Word Synset InductionabstractA word in natural language can be polysemous, having multiple meanings, as well as synonymous, meaning the same thing as other words. Word sense induction attempts to find the senses of polysemous words. Synonymy detection attempts to find when two words are interchangeable. We combine these tasks, first inducing word senses and then detecting similar senses to form word-sense synonym sets (synsets) in an unsupervised fashion. Given pairs of images and text with noun phrase labels, we perform synset induction to produce collections of underlying concepts described by one or more noun phrases. We find that considering multi-modal features from both visual and textual context yields better induced synsets than using either context alone. Human evaluations show that our unsupervised, multi-modally induced synsets are comparable in quality to annotation-assisted ImageNet synsets, achieving about 84% of ImageNet synsets' approval. Jesse Thomason, Raymond J. Mooney |
IJCAI | 1 |
| 2016 | Learning Multi-Modal Grounded Linguistic Semantics by Playing "I Spy"
Jesse Thomason, Jivko Sinapov, Maxwell Svetlik, Peter Stone 0001, Raymond J. Mooney |
IJCAI | 1 |
| 2015 | Learning to Interpret Natural Language Commands through Human-Robot Dialog
Jesse Thomason, Shiqi Zhang 0001, Raymond J. Mooney, Peter Stone 0001 |
IJCAI | 1 |
| 2014 | Integrating Language and Vision to Generate Natural Language Descriptions of Videos in the Wild
Jesse Thomason, Subhashini Venugopalan, Sergio Guadarrama, Kate Saenko, Raymond J. Mooney |
COLING | 1 |