Yoav Artzi

dblp:32/10489 · DBLP profile ↗
← Back
45ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0002-4605-6144ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Success and Cost Elicit Convention Formation for Efficient Communication
abstract
Humans leverage shared conversational context to become increasingly successful and efficient at communicating over time.One manifestation of this is the formation of ad hoc linguistic conventions, which allow people to coordinate on short, less costly utterances that are understood using shared conversational context.We present a method to train large multimodal models to form conventions, enabling efficient communication.Our approach uses simulated reference games between models, and requires no additional human-produced data.In repeated reference games involving photographs and tangram images, our method enables models to communicate efficiently with people: reducing the message length by up to 41% while increasing success by 15% over the course of the interaction.Human listeners respond faster when interacting with our model that forms conventions.We also show that training based on success or cost alone is insufficient -both are necessary to elicit convention formation.
Saujas Vaduguru, Yilun Hua, Yoav Artzi, Daniel Fried
ACL (1)3
2025 Retrospective Learning from Interactions
abstract
Multi-turn interactions between large language models (LLMs) and users naturally include implicit feedback signals. If an LLM responds in an unexpected way to an instruction, the user is likely to signal it by rephrasing the request, expressing frustration, or pivoting to an alternative task. Such signals are task-independent and occupy a relatively constrained subspace of language, allowing the LLM to identify them even if it fails on the actual task. We introduce ReSpect, a method to learn from such signals in past interactions via retrospection without additional annotations. We deploy ReSpect in a new multimodal interaction scenario, where humans instruct a multimodal LLM to solve an abstract reasoning task with a combinatorial solution space. Through thousands of interactions with humans, we show how ReSpect gradually improves task completion rate from 31% to 82%, all without any external annotation.
Zizhao Chen, Mustafa Omer Gul, Gloria Geng, Anne Wu, Yoav Artzi
ACL (1)6
2025 Imitation Learning from a Single Temporally Misaligned Video
abstract
We examine the problem of learning sequential tasks from a single visual demonstration. A key challenge arises when demonstrations are temporally misaligned due to variations in timing, differences in embodiment, or inconsistencies in execution. Existing approaches treat imitation as a distribution-matching problem, aligning individual frames between the agent and the demonstration. However, we show that such frame-level matching fails to enforce temporal ordering or ensure consistent progress. Our key insight is that matching should instead be defined at the level of sequences. We propose that perfect matching occurs when one sequence successfully covers all the subgoals in the same order as the other sequence. We present ORCA (ORdered Coverage Alignment), a dense per-timestep reward function that measures the probability of the agent covering demonstration frames in the correct order. On temporally misaligned demonstrations, we show that agents trained with the ORCA reward achieve $4.5$x improvement ($0.11 \rightarrow 0.50$ average normalized returns) for Meta-world tasks and $6.6$x improvement ($6.55 \rightarrow 43.3$ average returns) for Humanoid-v4 tasks compared to the best frame-level matching algorithms. We also provide empirical analysis showing that ORCA is robust to varying levels of temporal misalignment. The project website is at https://portal-cornell.github.io/orca/
William Huey, Huaxiaoyue Wang, Anne Wu, Yoav Artzi, Sanjiban Choudhury
ICML4
2025 Knot So Simple: A Minimalistic Environment for Spatial Reasoning
abstract
We propose KnotGym, an interactive environment for complex, spatial reasoning and manipulation. KnotGym includes goal-oriented rope manipulation tasks with varying levels of complexity, all requiring acting from pure image observations.Tasks are defined along a clear and quantifiable axis of complexity based on the number of knot crossings, creating a natural generalization test.KnotGym has a simple observation space, allowing for scalable development, yet it highlights core challenges in integrating acute perception, spatial reasoning, and grounded manipulation.We evaluate methods of different classes, including model-based RL, model-predictive control, and chain-of-thought reasoning, and illustrate the challenges KnotGym presents.
Zizhao Chen, Yoav Artzi
NeurIPS2
2024 CoGen: Learning from Feedback with Coupled Comprehension and Generation
abstract
Systems with both language comprehension and generation capabilities can benefit from the tight connection between the two.This work studies coupling comprehension and generation with focus on continually learning from interaction with users.We propose techniques to tightly integrate the two capabilities for both learning and inference.We situate our studies in two-player reference games, and deploy various models for thousands of interactions with human users, while learning from interaction feedback signals.We show dramatic improvements in performance over time, with comprehension-generation coupling leading to performance improvements up to 26% in absolute terms and up to 17% higher accuracies compared to a non-coupled system.Our analysis also shows coupling has substantial qualitative impact on the system's language, making it significantly more human-like.
Mustafa Omer Gul, Yoav Artzi
EMNLP2
2023 lilGym: Natural Language Visual Reasoning with Reinforcement Learning
abstract
We present lilGym, a new benchmark for language-conditioned reinforcement learning in visual environments.lilGym is based on 2,661 highly-compositional human-written natural language statements grounded in an interactive visual environment.We introduce a new approach for exact reward computation in every possible world state by annotating all statements with executable Python programs.Each statement is paired with multiple start states and reward functions to form thousands of distinct Markov Decision Processes of varying difficulty.We experiment with lilGym with different models and learning regimes.Our results and analysis show that while existing methods are able to achieve non-trivial performance, lilGym forms a challenging open problem.lilGym is available at https://lil.nlp.cornell.edu/lilgym/.
Anne Wu, Kianté Brantley, Noriyuki Kojima, Yoav Artzi
ACL (1)4
2023 Semantic uncertainty guides the extension of conventions to new referents
Ron Eliav, Anya Ji, Yoav Artzi, Robert D. Hawkins
CogSci3
2023 Continually Improving Extractive QA via Human Feedback
abstract
We study continually improving an extractive question answering (QA) system via human user feedback.We design and deploy an iterative approach, where information-seeking users ask questions, receive model-predicted answers, and provide feedback.We conduct experiments involving thousands of user interactions under diverse setups to broaden the understanding of learning from feedback over time.Our experiments show effective improvement from user feedback of extractive QA models over time across different data regimes, including significant potential for domain adaptation.* Equal contribution. 1 The term continual learning is at times used to refer to a scenario where models adapt to new tasks over time.We study improving the model continually on its original task.
Hung-Ting Chen, Yoav Artzi, Eunsol Choi
EMNLP3
2023 Wav2Seq: Pre-Training Speech-to-Text Encoder-Decoder Models Using Pseudo Languages
abstract
We introduce Wav2Seq, the first self-supervised approach to pre-train both parts of encoder-decoder models for speech data. We induce a pseudo language as a compact discrete representation, and formulate a self-supervised pseudo speech recognition task — transcribing audio inputs into pseudo subword sequences. This process stands on its own, or can be applied as low-cost second-stage pre-training. We experiment with automatic speech recognition (ASR), spoken named entity recognition, and speech-to-text translation. We set new state-of-the-art results for end-to-end spoken named entity recognition, and show consistent improvements on 8 language pairs for speech-to-text translation, even when competing methods use additional text data for training. On ASR, our approach enables encoder-decoder methods to benefit from pre-training for all parts of the network, and shows comparable performance to highly optimized recent methods.
Felix Wu, Kwangyoun Kim, Shinji Watanabe 0001, Kyu Jeong Han, Ryan McDonald, Kilian Q. Weinberger, Yoav Artzi
ICASSP7
2023 IncDSI: Incrementally Updatable Document Retrieval
abstract
Differentiable Search Index is a recently proposed paradigm for document retrieval, that encodes information about a corpus of documents within the parameters of a neural network and directly maps queries to corresponding documents. These models have achieved state-of-the-art performances for document retrieval across many benchmarks. These kinds of models have a significant limitation: it is not easy to add new documents after a model is trained. We propose IncDSI, a method to add documents in real time (about 20-50ms per document), without retraining the model on the entire dataset (or even parts thereof). Instead we formulate the addition of documents as a constrained optimization problem that makes minimal changes to the network parameters. Although orders of magnitude faster, our approach is competitive with re-training the model on the whole dataset and enables the development of document retrieval systems that can be updated with new information in real-time. Our code for IncDSI is available at https://github.com/varshakishore/IncDSI.
Varsha Kishore, Chao Wan, Justin Lovelace, Yoav Artzi, Kilian Q. Weinberger
ICML4
2023 Continual Learning for Instruction Following from Realtime Feedback
abstract
We propose and deploy an approach to continually train an instruction-following agent from feedback provided by users during collaborative interactions. During interaction, human users instruct an agent using natural language, and provide realtime binary feedback as they observe the agent following their instructions. We design a contextual bandit learning approach, converting user feedback to immediate reward. We evaluate through thousands of human-agent interactions, demonstrating 15.4% absolute improvement in instruction execution accuracy over time. We also show our approach is robust to several design variations, and that the feedback signal is roughly equivalent to the learning signal of supervised demonstration data.
Alane Suhr, Yoav Artzi
NeurIPS2
2022 Simulating Bandit Learning from User Feedback for Extractive Question Answering
abstract
We study learning from user feedback for extractive question answering by simulating feedback using supervised data.We cast the problem as contextual bandit learning, and analyze the characteristics of several learning scenarios with focus on reducing data annotation.We show that systems initially trained on a small number of examples can dramatically improve given feedback from users on modelpredicted answers, and that one can use existing datasets to deploy systems in new domains without any annotation, but instead improving the system on-the-fly via user feedback.
Eunsol Choi, Yoav Artzi
ACL (1)3
2022 Abstract Visual Reasoning with Tangram Shapes
abstract
We introduce KILOGRAM, a resource for studying abstract visual reasoning in humans and machines.Drawing on the history of tangram puzzles as stimuli in cognitive science, we build a richly annotated dataset that, with >1k distinct stimuli, is orders of magnitude larger and more diverse than prior resources.It is both visually and linguistically richer, moving beyond whole shape descriptions to include segmentation maps and part labels.We use this resource to evaluate the abstract visual reasoning capacities of recent multi-modal models.We observe that pre-trained weights demonstrate limited abstract reasoning, which dramatically improves with fine-tuning.We also observe that explicitly describing parts aids abstract reasoning for both humans and models, especially when jointly encoding the linguistic and visual inputs.
Anya Ji, Noriyuki Kojima, Noah Rush, Alane Suhr, Wai Keen Vong, Robert D. Hawkins, Yoav Artzi
EMNLP7
2022 SLUE: New Benchmark Tasks For Spoken Language Understanding Evaluation on Natural Speech
abstract
Progress in speech processing has been facilitated by shared datasets and benchmarks. Historically these have focused on automatic speech recognition (ASR), speaker identification, or other lower-level tasks. Interest has been growing in higher-level spoken language understanding tasks, including using end-to-end models, but there are fewer annotated datasets for such tasks. At the same time, recent work shows the possibility of pre-training generic representations and then fine-tuning for several tasks using relatively little labeled data. We propose to create a suite of benchmark tasks for Spoken Language Understanding Evaluation (SLUE) consisting of limited-size labeled training sets and corresponding evaluation sets. This resource would allow the research community to track progress, evaluate pre-trained representations for higher-level tasks, and study open questions such as the utility of pipeline versus end-to-end approaches. We present the first phase of the SLUE benchmark suite, consisting of named entity recognition, sentiment analysis, and ASR on the corresponding datasets. We focus on naturally produced (not read or synthesized) speech, and freely available datasets. We pro-vide new transcriptions and annotations on subsets of the VoxCeleb and VoxPopuli datasets, evaluation metrics and results for baseline models, and an open-source toolkit to reproduce the baselines and evaluate new models.
Suwon Shon, Ankita Pasad, Felix Wu, Pablo Brusco, Yoav Artzi, Karen Livescu, Kyu Jeong Han
ICASSP5
2022 Performance-Efficiency Trade-Offs in Unsupervised Pre-Training for Speech Recognition
abstract
This paper is a study of performance-efficiency trade-offs in pre-trained models for automatic speech recognition (ASR). We focus on wav2vec 2.0, and formalize several architecture designs that influence both the model performance and its efficiency. Putting together all our observations, we introduce SEW-D (Squeezed and Efficient Wav2vec with Disentangled Attention), a pre-trained model architecture with significant improvements along both performance and efficiency dimensions across a variety of training setups. For example, under the 100h-960h semi-supervised setup on LibriSpeech, SEW-D achieves a 1.9x inference speedup compared to wav2vec 2.0, with a 13.5% relative reduction in word error rate. With a similar inference time, SEW reduces word error rate by 25–50% across different model sizes.
Felix Wu, Kwangyoun Kim, Kyu Jeong Han, Kilian Q. Weinberger, Yoav Artzi
ICASSP6
2022 Spoken language interaction with robots: Recommendations for future research
abstract
With robotics rapidly advancing, more effective human–robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language interaction capabilities is still very limited. In this article, based on the report of an interdisciplinary workshop convened by the National Science Foundation, we identify key scientific and engineering advances needed to enable effective spoken language interaction with robotics. We make 25 recommendations, involving eight general themes: putting human needs first, better modeling the social and interactive aspects of language, improving robustness, creating new methods for rapid adaptation, better integrating speech and language with other communication modalities, giving speech and language components access to rich representations of the robot’s current knowledge and state, making all components operate in real time, and improving research infrastructure and resources. Research and development that prioritizes these topics will, we believe, provide a solid foundation for the creation of speech-capable robots that are easy and effective for humans to work with.
Matthew Marge, Carol Y. Espy-Wilson, Nigel G. Ward, Abeer Alwan, Yoav Artzi, Mohit Bansal, Gilmer L. Blankenship, Joyce Y. Chai, Hal Daumé III, Debadeepta Dey, Mary P. Harper, Thomas Howard, Casey Kennington, Ivana Kruijff-Korbayová, Dinesh Manocha, Cynthia Matuszek, Ross Mead, Raymond J. Mooney, Roger K. Moore, Mari Ostendorf, Heather Pon-Barry, Alexander I. Rudnicky, Matthias Scheutz, Robert St. Amant, Stefanie Tellex, David R. Traum, Zhou Yu 0005
Comput. Speech Lang.5
2021 Who's Waldo? Linking People Across Text and Images
abstract
We present a task and benchmark dataset for person-centric visual grounding, the problem of linking between people named in a caption and people pictured in an image. In contrast to prior work in visual grounding, which is predominantly object-based, our new task masks out the names of people in captions in order to encourage methods trained on such image–caption pairs to focus on contextual cues, such as the rich interactions between multiple people, rather than learning associations between names and appearances. To facilitate this task, we introduce a new dataset, Who’s Waldo, mined automatically from image–caption data on Wikimedia Commons. We propose a Transformer-based method that outperforms several strong baselines on this task, and release our data to the research community to spur work on contextual models that consider both vision and language. Code and data are available at: https://whoswaldo.github.io
Claire Yuqing Cui, Apoorv Khandelwal 0001, Yoav Artzi, Noah Snavely, Hadar Averbuch-Elor
ICCV3
2021 Revisiting Few-sample BERT Fine-tuning
Tianyi Zhang 0007, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger, Yoav Artzi
ICLR5
2021 Continual Learning for Grounded Instruction Generation by Observing Human Following Behavior
abstract
Abstract We study continual learning for natural language instruction generation, by observing human users’ instruction execution. We focus on a collaborative scenario, where the system both acts and delegates tasks to human users using natural language. We compare user execution of generated instructions to the original system intent as an indication to the system’s success communicating its intent. We show how to use this signal to improve the system’s ability to generate instructions via contextual bandit learning. In interaction with real users, our system demonstrates dramatic improvements in its ability to generate language over time.
Noriyuki Kojima, Alane Suhr, Yoav Artzi
Trans. Assoc. Comput. Linguistics3
2020 What is Learned in Visually Grounded Neural Syntax Acquisition
abstract
Visual features are a promising signal for learning bootstrap textual models.However, blackbox learning models make it difficult to isolate the specific contribution of visual components.In this analysis, we consider the case study of the Visually Grounded Neural Syntax Learner (Shi et al., 2019), a recent approach for learning syntax from a visual training signal.By constructing simplified versions of the model, we isolate the core factors that yield the model's strong performance.Contrary to what the model might be capable of learning, we find significantly less expressive versions produce similar predictions and perform just as well, or even better.We also find that a simple lexical signal of noun concreteness plays the main role in the model's predictions as opposed to more complex syntactic reasoning.
Noriyuki Kojima, Hadar Averbuch-Elor, Alexander M. Rush, Yoav Artzi
ACL4
2020 Interactive Classification by Asking Informative Questions
abstract
We study the potential for interaction in natural language classification.We add a limited form of interaction for intent classification, where users provide an initial query using natural language, and the system asks for additional information using binary or multichoice questions.At each turn, our system decides between asking the most informative question or making the final classification prediction.The simplicity of the model allows for bootstrapping of the system without interaction data, instead relying on simple crowdsourcing tasks.We evaluate our approach on two domains, showing the benefit of interaction and the advantage of learning to balance between asking additional questions and making the final prediction.What is the bill length of the bird: shorter, similar, or longer than head?Shorter than head.Is the bird underpart orange?Yes.The identified bird is: American Redstart FAQ Suggestion What data limits apply when roaming internationally?American Crow Bobolink … American Redstart How do I sign up for Sprint Global Roaming? . . .How do I purchase a High Speed Data Roaming Pass?Bird Identification Travel out of country.Do you need to activate global roaming service?Yes.Do you want high speed data roaming?No.
Lili Yu, Howard Chen 0003, Sida I. Wang, Tao Lei 0001, Yoav Artzi
ACL5
2020 BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang 0007, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, Yoav Artzi
ICLR5
2019 A Corpus for Reasoning about Natural Language Grounded in Photographs
abstract
We introduce a new dataset for joint reasoning about natural language and images, with a focus on semantic diversity, compositionality, and visual reasoning challenges.The data contains 107,292 examples of English sentences paired with web photographs.The task is to determine whether a natural language caption is true about a pair of photographs.We crowdsource the data using sets of visually rich images and a compare-and-contrast task to elicit linguistically diverse language.Qualitative analysis shows the data requires compositional joint reasoning, including about quantities, comparisons, and relations.Evaluation using state-of-the-art visual reasoning methods shows the data presents a strong challenge.* Contributed equally.† Work done as an undergraduate at Cornell University. 1 In parts of this paper, we use the term compositional differently than it is commonly used in linguistics to refer to reasoning that requires composition.This type of reasoning often manifests itself in highly compositional language.
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, Yoav Artzi
ACL (1)6
2019 TOUCHDOWN: Natural Language Navigation and Spatial Reasoning in Visual Street Environments
abstract
We study the problem of jointly reasoning about language and vision through a navigation and spatial reasoning task. We introduce the Touchdown task and dataset, where an agent must first follow navigation instructions in a Street View environment to a goal position, and then guess a location in its observed environment described in natural language to find a hidden object. The data contains 9326 examples of English instructions and spatial descriptions paired with demonstrations. We perform qualitative linguistic analysis, and show that the data displays a rich use of spatial reasoning. Empirical analysis shows the data presents an open challenge to existing methods.
Howard Chen 0003, Alane Suhr, Dipendra Misra, Noah Snavely, Yoav Artzi
CVPR5
2019 Executing Instructions in Situated Collaborative Interactions
abstract
Alane Suhr, Claudia Yan, Jack Schluger, Stanley Yu, Hadi Khader, Marwa Mouallem, Iris Zhang, Yoav Artzi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Alane Suhr, Claudia Yan, Jacob Schluger, Stanley Yu, Hadi Khader, Marwa Mouallem, Iris Zhang, Yoav Artzi
EMNLP/IJCNLP (1)8
2019 EARLY FUSION for Goal Directed Robotic Vision
abstract
Building perceptual systems for robotics which perform well under tight computational budgets requires novel architectures which rethink the traditional computer vision pipeline. Modern vision architectures require the agent to build a summary representation of the entire scene, even if most of the input is irrelevant to the agent's current goal. In this work, we flip this paradigm, by introducing EARLYFUSION vision models that condition on a goal to build custom representations for downstream tasks. We show that these goal specific representations can be learned more quickly, are substantially more parameter efficient, and more robust than existing attention mechanisms in our domain. We demonstrate the effectiveness of these methods on a simulated item retrieval problem that is trained in a fully end-to-end manner via imitation learning.
Aaron Walsman, Yonatan Bisk, Saadia Gabriel, Dipendra Misra, Yoav Artzi, Yejin Choi 0001, Dieter Fox
IROS5
2019 Understanding Reader Backtracking Behavior in Online News Articles
abstract
Rich engagement data can shed light on how people interact with online content and how such interactions may be determined by the content of the page. In this work, we investigate a specific type of interaction, backtracking, which refers to the action of scrolling back in a browser while reading an online news article. We leverage a dataset of close to 700K instances of more than 15K readers interacting with online news articles, in order to characterize and predict backtracking behavior. We first define different types of backtracking actions. We then show that “full” backtracks, where the readers eventually return to the spot at which they left the text, can be predicted by using features that were previously shown to relate to text readability. This finding highlights the relationship between backtracking and readability and suggests that backtracking could help assess readability of content at scale.
Uzi Smadja, Max Grusky, Yoav Artzi, Mor Naaman
WWW3
2018 Situated Mapping of Sequential Instructions to Actions with Single-step Reward Observation
abstract
We propose a learning approach for mapping context-dependent sequential instructions to actions.We address the problem of discourse and state dependencies with an attention-based model that considers both the history of the interaction and the state of the world.To train from start and goal states without access to demonstrations, we propose SESTRA, a learning algorithm that takes advantage of singlestep reward observations and immediate expected reward maximization.We evaluate on the SCONE domains, and show absolute accuracy improvements of 9.8%-25.3%across the domains over approaches that use high-level logical representations.
Alane Suhr, Yoav Artzi
ACL (1)2
2018 Simple Recurrent Units for Highly Parallelizable Recurrence
abstract
Common recurrent neural architectures scale poorly due to the intrinsic difficulty in parallelizing their state computations.In this work, we propose the Simple Recurrent Unit (SRU), a light recurrent unit that balances model capacity and scalability.SRU is designed to provide expressive recurrence, enable highly parallelized implementation, and comes with careful initialization to facilitate training of deep models.We demonstrate the effectiveness of SRU on multiple NLP tasks.SRU achieves 5-9x speed-up over cuDNN-optimized LSTM on classification and question answering datasets, and delivers stronger results than LSTM and convolutional models.We also obtain an average of 0.7 BLEU improvement over the Transformer model (Vaswani et al., 2017) on translation by incorporating SRU into the architecture.1
Tao Lei 0001, Yu Zhang 0033, Sida I. Wang, Hui Dai, Yoav Artzi
EMNLP5
2018 Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction
abstract
We propose to decompose instruction execution to goal prediction and action generation.We design a model that maps raw visual observations to goals using LINGUNET, a language-conditioned image generation network, and then generates the actions required to complete them.Our model is trained from demonstration only without external resources.To evaluate our approach, we introduce two benchmarks for instruction following: LANI, a navigation task; and CHAI, where an agent executes household instructions.Our evaluation demonstrates the advantages of our model decomposition, and illustrates the challenges posed by our new benchmarks.
Dipendra Misra, Andrew Bennett, Valts Blukis, Eyvind Niklasson, Max Shatkhin, Yoav Artzi
EMNLP6
2018 Newsroom: A Dataset of 1.3 Million Summaries with Diverse Extractive Strategies
abstract
Max Grusky, Mor Naaman, Yoav Artzi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Max Grusky, Mor Naaman, Yoav Artzi
NAACL-HLT3
2018 Learning to Map Context-Dependent Sentences to Executable Formal Queries
abstract
Alane Suhr, Srinivasan Iyer, Yoav Artzi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Alane Suhr, Srinivasan Iyer 0001, Yoav Artzi
NAACL-HLT3
2017 Modeling Sub-Document Attention Using Viewport Time
abstract
Website measures of engagement captured from millions of users, such as in-page scrolling and viewport position, can provide deeper understanding of attention than possible with simpler measures, such as dwell time. Using data from 1.2M news reading sessions, we examine and evaluate three increasingly sophisticated models of sub-document attention computed from viewport time, the time a page component is visible on the user display. Our modeling incorporates prior eye-tracking knowledge about onscreen reading, and we validate it by showing how, when used to estimate user reading rate, it aligns with known empirical measures. We then show how our models reveal an interaction between article topic and attention to page elements. Our approach supports refined large-scale measurement of user engagement at a level previously available only from lab-based eye-tracking studies.
Max Grusky, Jeiran Jahani, Josh Schwartz, Dan Valente, Yoav Artzi, Mor Naaman
CHI5
2017 Mapping Instructions and Visual Observations to Actions with Reinforcement Learning
abstract
We propose to directly map raw visual observations and text input to actions for instruction execution.While existing approaches assume access to structured environment representations or use a pipeline of separately trained models, we learn a single model to jointly reason about linguistic and visual input.We use reinforcement learning in a contextual bandit setting to train a neural network agent.To guide the agent's exploration, we use reward shaping with different forms of supervision.Our approach does not require intermediate representations, planning procedures, or training different models.We evaluate in a simulated environment, and show significant improvements over supervised learning and common reinforcement learning variants.
Dipendra Misra, John Langford 0001, Yoav Artzi
EMNLP3
2016 Neural Shift-Reduce CCG Semantic Parsing
Dipendra Misra, Yoav Artzi
EMNLP2
2015 Broad-coverage CCG Semantic Parsing with AMR
abstract
We propose a grammar induction technique for AMR semantic parsing.While previous grammar induction techniques were designed to re-learn a new parser for each target application, the recently annotated AMR Bank provides a unique opportunity to induce a single model for understanding broad-coverage newswire text and support a wide range of applications.We present a new model that combines CCG parsing to recover compositional aspects of meaning and a factor graph to model non-compositional phenomena, such as anaphoric dependencies.Our approach achieves 66.2 Smatch F1 score on the AMR bank, significantly outperforming the previous state of the art.
Yoav Artzi, Kenton Lee, Luke Zettlemoyer
EMNLP1
2015 Event Detection and Factuality Assessment with Non-Expert Supervision
abstract
Events are communicated in natural language with varying degrees of certainty.For example, if you are "hoping for a raise," it may be somewhat less likely than if you are "expecting" one.To study these distinctions, we present scalable, highquality annotation schemes for event detection and fine-grained factuality assessment.We find that non-experts, with very little training, can reliably provide judgments about what events are mentioned and the extent to which the author thinks they actually happened.We also show how such data enables the development of regression models for fine-grained scalar factuality predictions that outperform strong baselines.
Kenton Lee, Yoav Artzi, Yejin Choi 0001, Luke Zettlemoyer
EMNLP2
2014 Learning to Automatically Solve Algebra Word Problems
abstract
We present an approach for automatically learning to solve algebra word problems.Our algorithm reasons across sentence boundaries to construct and solve a system of linear equations, while simultaneously recovering an alignment of the variables and numbers in these equations to the problem text.The learning algorithm uses varied supervision, including either full equations or just the final answers.We evaluate performance on a newly gathered corpus of algebra word problems, demonstrating that the system can correctly answer almost 70% of the questions in the dataset.This is, to our knowledge, the first learning result for this task.
Nate Kushman, Luke Zettlemoyer, Regina Barzilay, Yoav Artzi
ACL (1)4
2014 Context-dependent Semantic Parsing for Time Expressions
abstract
We present an approach for learning context-dependent semantic parsers to identify and interpret time expressions. We use a Combinatory Categorial Grammar to construct compositional meaning representations, while considering contextual cues, such as the document creation time and the tense of the governing verb, to compute the final time values. Experiments on benchmark datasets show that our approach outperforms previous stateof-the-art systems, with error reductions of 13% to 21% in end-to-end performance.
Kenton Lee, Yoav Artzi, Jesse Dodge, Luke Zettlemoyer
ACL (1)2
2014 Learning Compact Lexicons for CCG Semantic Parsing
abstract
We present methods to control the lexicon size when learning a Combinatory Categorial Grammar semantic parser.Existing methods incrementally expand the lexicon by greedily adding entries, considering a single training datapoint at a time.We propose using corpus-level statistics for lexicon learning decisions.We introduce voting to globally consider adding entries to the lexicon, and pruning to remove entries no longer required to explain the training data.Our methods result in state-of-the-art performance on the task of executing sequences of natural language instructions, achieving up to 25% error reduction, with lexicons that are up to 70% smaller and are qualitatively less noisy.* This research was carried out at Google.
Yoav Artzi, Dipanjan Das 0001, Slav Petrov
EMNLP1
2013 Learning Distributions over Logical Forms for Referring Expression Generation
abstract
We present a new approach to referring expression generation, casting it as a density estimation problem where the goal is to learn distributions over logical expressions identifying sets of objects in the world.Despite an extremely large space of possible expressions, we demonstrate effective learning of a globally normalized log-linear distribution.This learning is enabled by a new, multi-stage approximate inference technique that uses a pruning model to construct only the most likely logical forms.We train and evaluate the approach on a new corpus of references to sets of visual objects.Experiments show the approach is able to learn accurate models, which generate over 87% of the expressions people used.Additionally, on the previously studied special case of single object reference, we show a 35% relative error reduction over previous state of the art.
Nicholas FitzGerald, Yoav Artzi, Luke Zettlemoyer
EMNLP2
2013 Scaling Semantic Parsers with On-the-Fly Ontology Matching
abstract
We consider the challenge of learning semantic parsers that scale to large, open-domain problems, such as question answering with Freebase.In such settings, the sentences cover a wide variety of topics and include many phrases whose meaning is difficult to represent in a fixed target ontology.For example, even simple phrases such as 'daughter' and 'number of people living in' cannot be directly represented in Freebase, whose ontology instead encodes facts about gender, parenthood, and population.In this paper, we introduce a new semantic parsing approach that learns to resolve such ontological mismatches.The parser is learned from question-answer pairs, uses a probabilistic CCG to build linguistically motivated logicalform meaning representations, and includes an ontology matching model that adapts the output logical forms for each target ontology.Experiments demonstrate state-of-the-art performance on two benchmark semantic parsing datasets, including a nine point accuracy improvement on a recent Freebase QA corpus.
Tom Kwiatkowski, Eunsol Choi, Yoav Artzi, Luke Zettlemoyer
EMNLP3
2013 Weakly Supervised Learning of Semantic Parsers for Mapping Instructions to Actions
abstract
The context in which language is used provides a strong signal for learning to recover its meaning. In this paper, we show it can be used within a grounded CCG semantic parsing approach that learns a joint model of meaning and context for interpreting and executing natural language instructions, using various types of weak supervision. The joint nature provides crucial benefits by allowing situated cues, such as the set of visible objects, to directly influence learning. It also enables algorithms that learn while executing instructions, for example by trying to replicate human actions. Experiments on a benchmark navigational dataset demonstrate strong performance under differing forms of supervision, including correctly executing 60% more instruction sets relative to the previous state of the art.
Yoav Artzi, Luke Zettlemoyer
Trans. Assoc. Comput. Linguistics1
2012 Predicting Responses to Microblog Posts
Yoav Artzi, Patrick Pantel, Michael Gamon
HLT-NAACL1
2011 Bootstrapping Semantic Parsers from Conversations
Yoav Artzi, Luke Zettlemoyer
EMNLP1