Allyson Ettinger

dblp:165/0758 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
16since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 4 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 To Err Is AI: A Case Study Informing LLM Flaw Reporting Practices
abstract
In August of 2024, 495 hackers generated evaluations in an open-ended bug bounty targeting the Open Language Model (OLMo) from The Allen Institute for AI. A vendor panel staffed by representatives of OLMo's safety program adjudicated changes to OLMo's documentation and awarded cash bounties to participants who successfully demonstrated a need for public disclosure clarifying the intent, capacities, and hazards of model deployment. This paper presents a collection of lessons learned, illustrative of flaw reporting best practices intended to reduce the likelihood of incidents and produce safer large language models (LLMs). These include best practices for safety reporting processes, their artifacts, and safety program staffing.
Sean McGregor, Allyson Ettinger, Nick Judd, Paul Albee, Kavel Rao, William H. Smith, Shayne Longpre, Avijit Ghosh, Christopher Fiorelli, Michelle Hoang, Sven Cattell, Nouha Dziri
AAAI2
2025 AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
abstract
Creativity has long been considered one of the most difficult aspect of human intelligence for AI to mimic. However, the rise of Large Language Models (LLMs), like ChatGPT, has raised questions about whether AI can match or even surpass human creativity. We present CREATIVITY INDEX as the first step to quantify the linguistic creativity of a text by reconstructing it from existing text snippets on the web. CREATIVITY INDEX is motivated by the hypothesis that the seemingly remarkable creativity of LLMs may be attributable in large part to the creativity of human-written texts on the web. To compute CREATIVITY INDEX efficiently, we introduce DJ SEARCH, a novel dynamic programming algorithm that can search verbatim and near-verbatim matches of text snippets from a given document against the web. Experiments reveal that the CREATIVITY INDEX of professional human authors is on average 66.2% higher than that of LLMs, and that alignment reduces the CREATIVITY INDEX of LLMs by an average of 30.1%. In addition, we find that distinguished authors like Hemingway exhibit measurably higher CREATIVITY INDEX compared to other human writers. Finally, we demonstrate that CREATIVITY INDEX can be used as a surprisingly effective criterion for zero-shot machine text detection, surpassing the strongest existing zero-shot system, DetectGPT, by a significant margin of 30.2%, and even outperforming the strongest supervised system, GhostBuster, in five out of six domains.
Ximing Lu, Melanie Sclar, Skyler Hallinan, Niloofar Mireshghallah, Jiacheng Liu 0010, Seungju Han 0002, Allyson Ettinger, Khyathi Raghavi Chandu, Nouha Dziri, Yejin Choi 0001
ICLR7
2025 FlexOLMo: Open Language Models for Flexible Data Use
abstract
We introduce FlexOLMo, a new class of language models (LMs) that supports (1) distributed training without data sharing, where different model parameters are independently trained on private datasets, and (2) data-flexible inference, where these parameters along with their associated data can be easily included or excluded from model inferences with no further training. FlexOLMo employs a mixture-of-experts (MoE) architecture where each expert is trained independently on private datasets and later integrated through a new nonparametric routing without any joint training across datasets. FlexOLMo is trained on FLEXMIX, a corpus we curate comprising seven restricted sets, either real or realistic approximations, alongside publicly available datasets. We evaluate models with up to 37 billion parameters (20 billion active) on 31 diverse downstream tasks. We show that a general expert trained on public data can be effectively combined with independently trained experts from other data owners significantly benefiting from these restricted sets (an average 41% relative improvement) while allowing flexible opt-out at inference time (e.g., for users without appropriate licenses or permissions). Our approach also outperforms prior model merging methods by 10.1% on average and surpasses the standard MoE trained without data restrictions using the same training FLOPs. Altogether, FlexOLMo enables training on restricted data while keeping data local and supports fine-grained control of data access at inference.
Akshita Bhagia, Kevin Farhat, Niklas Muennighoff, Jacob Morrison, Pete Walsh 0001, Dustin Schwenk, Shayne Longpre, Jake Poznanski, Allyson Ettinger, Daogao Liu, Margaret Li, Mike Lewis, Scott Yih, Dirk Groeneveld, Luca Soldaini, Kyle Lo, Noah A. Smith, Luke Zettlemoyer, Pang Wei W. Koh, Hannaneh Hajishirzi, Ali Farhadi, Sewon Min
NeurIPS10
2024 Experimental Contexts Can Facilitate Robust Semantic Property Inference in Language Models, but Inconsistently
abstract
Recent zero-shot evaluations have highlighted important limitations in the abilities of language models (LMs) to perform meaning extraction.However, it is now well known that LMs can demonstrate radical improvements in the presence of experimental contexts such as in-context examples and instructions.How well does this translate to previously studied meaning-sensitive tasks?We present a casestudy on the extent to which experimental contexts can improve LMs' robustness in performing property inheritance-predicting semantic properties of novel concepts, a task that they have been previously shown to fail on.Upon carefully controlling the nature of the in-context examples and the instructions, our work reveals that they can indeed lead to nontrivial property inheritance behavior in LMs.However, this ability is inconsistent: with a minimal reformulation of the task, some LMs were found to pick up on shallow, non-semantic heuristics from their inputs, suggesting that the computational principles of semantic property inference are yet to be mastered by LMs.
Kanishka Misra, Allyson Ettinger, Kyle Mahowald
EMNLP2
2024 The Generative AI Paradox: "What It Can Create, It May Not Understand"
abstract
The recent wave of generative AI has sparked unprecedented global attention, with both excitement and concern over potentially superhuman levels of artificial intelligence: models now take only seconds to produce outputs that would challenge or exceed the capabilities even of expert humans. At the same time, models still show basic errors in understanding that would not be expected even in non-expert humans. This presents us with an apparent paradox: how do we reconcile seemingly superhuman capabilities with the persistence of errors that few humans would make? In this work, we posit that this tension reflects a divergence in the configuration of intelligence in today's generative models relative to intelligence in humans. Specifically, we propose and test the **Generative AI Paradox** hypothesis: generative models, having been trained directly to reproduce expert-like outputs, acquire generative capabilities that are not contingent upon---and can therefore exceed---their ability to understand those same types of outputs. This contrasts with humans, for whom basic understanding almost always precedes the ability to generate expert-level outputs. We test this hypothesis through controlled experiments analyzing generation vs.~understanding in generative models, across both language and image modalities. Our results show that although models can outperform humans in generation, they consistently fall short of human capabilities in measures of understanding, as well as weaker correlation between generation and understanding performance, and more brittleness to adversarial inputs. Our findings support the hypothesis that models' generative capability may not be contingent upon understanding capability, and call for caution in interpreting artificial intelligence by analogy to human intelligence.
Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman, Jena D. Hwang, Jillian Fisher, Abhilasha Ravichander, Khyathi Raghavi Chandu, Benjamin Newman, Pang Wei Koh, Allyson Ettinger, Yejin Choi 0001
ICLR13
2024 WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
abstract
We introduce WildGuard---an open, light-weight moderation tool for LLM safety that achieves three goals: (1) identifying malicious intent in user prompts, (2) detecting safety risks of model responses, and (3) determining model refusal rate. Together, WildGuard serves the increasing needs for automatic safety moderation and evaluation of LLM interactions, providing a one-stop tool with enhanced accuracy and broad coverage across 13 risk categories. While existing open moderation tools such as Llama-Guard2 score reasonably well in classifying straightforward model interactions, they lag far behind a prompted GPT-4, especially in identifying adversarial jailbreaks and in evaluating models' refusals, a key measure for evaluating safety behaviors in model responses. To address these challenges, we construct WildGuardMix, a large-scale and carefully balanced multi-task safety moderation dataset with 92K labeled examples that cover vanilla (direct) prompts and adversarial jailbreaks, paired with various refusal and compliance responses. WildGuardMix is a combination of WildGuardTrain, the training data of WildGuard, and WildGuardTest, a high-quality human-annotated moderation test set with 5K labeled items covering broad risk scenarios.Through extensive evaluations on WildGuardTest and ten existing public benchmarks, we show that WildGuard establishes state-of-the-art performance in open-source safety moderation across all the three tasks compared to ten strong existing open-source moderation models (e.g., up to 25.3% improvement on refusal detection). Importantly, WildGuard matches and sometimes exceeds GPT-4 performance (e.g., up to 4.8% improvement on prompt harmfulness identification). WildGuard serves as a highly effective safety moderator in an LLM interface, reducing the success rate of jailbreak attacks from 79.8% to 2.4%. We will make all our data, models and training/evaluation code publicly available under CC BY 4.0 license.
Seungju Han 0002, Kavel Rao, Allyson Ettinger, Bill Y. Lin, Nathan Lambert 0001, Yejin Choi 0001, Nouha Dziri
NeurIPS3
2024 WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
abstract
We introduce WildTeaming, an automatic red-teaming framework that mines in-the-wild user-chatbot interactions to discover 5.7K unique clusters of novel jailbreak tactics, and then composes selections of multiple mined tactics for systematic exploration of novel and even more challenging jailbreaks. Compared to prior work that performed red-teaming via recruited human workers, gradient-based optimization, or iterative revision with large language models (LLMs), our work investigates jailbreaks from chatbot users in-the-wild who were not specifically instructed to break the system. WildTeaming reveals previously unidentified vulnerabilities of frontier LLMs, resulting in more diverse and successful adversarial attacks compared to state-of-the-art jailbreaking methods. While there exist many datasets for jailbreak evaluation, very few open-source datasets exist for jailbreak training, as safety training data has been closed among all frontier models even when their weights are open. Therefore, with WildTeaming we create WildJailbreak, a large-scale open-source synthetic safety dataset with 262K vanilla (direct request) and adversarial (complex jailbreak) prompt-response pairs. In order to mitigate exaggerated safety behaviors, WildJailbreak provides two contrastive types of queries: 1) harmful queries (both vanilla and adversarial) and 2) benign queries that resemble harmful queries in form but contain no harmful intent. As WildJailbreak considerably upgrades the quality and scale of existing safety resources, it uniquely enables us to examine the scaling effects of data and the interplay of data properties and model capabilities during safety training. Through extensive model training and evaluations, we identify the training properties that enable an ideal balance of safety behaviors: appropriate safeguarding without over-refusal, effective handling of both vanilla and adversarial queries, and minimal, if any, decrease in general capabilities. All the components of WildJailbreak contribute to achieving balanced safety behaviors of models
Kavel Rao, Seungju Han 0002, Allyson Ettinger, Faeze Brahman, Sachin Kumar 0009, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi 0001, Nouha Dziri
NeurIPS4
2023 COMPS: Conceptual Minimal Pair Sentences for testing Robust Property Knowledge and its Inheritance in Pre-trained Language Models
abstract
A characteristic feature of human semantic cognition is its ability to not only store and retrieve the properties of concepts observed through experience, but to also facilitate the inheritance of properties (can breathe) from superordinate concepts (ANIMAL) to their subordinates (DOG)-i.e.demonstrate property inheritance.In this paper, we present COMPS, a collection of English minimal pair sentences that jointly tests pre-trained language models (PLMs) on their ability to attribute properties to concepts and their ability to demonstrate property inheritance behavior.Analyses of 22 different PLMs on COMPS reveal that they can easily distinguish between concepts on the basis of a property when they are trivially different, but find it relatively difficult when concepts are related on the basis of nuanced knowledge representations.Furthermore, we find that PLMs can show behaviors suggesting successful property inheritance in simple contexts, but fail in the presence of distracting information, which decreases the performance of many models sometimes even below chance.This lack of robustness in demonstrating simple reasoning raises important questions about PLMs' capacity to make correct inferences even when they appear to possess the prerequisite knowledge.
Kanishka Misra, Julia Taylor Rayz, Allyson Ettinger
EACL3
2023 Can You Follow Me? Testing Situational Understanding for ChatGPT
abstract
Understanding sentence meanings and updating information states appropriately across time-what we call "situational understanding" (SU)-is a critical ability for human-like AI agents.SU is essential in particular for chat models, such as ChatGPT, to enable consistent, coherent, and effective dialogue between humans and AI.Previous works have identified certain SU limitations in non-chatbot Large Language models (LLMs), but the extent and causes of these limitations are not well understood, and capabilities of current chat-based models in this domain have not been explored.In this work we tackle these questions, proposing a novel synthetic environment for SU testing which allows us to do controlled and systematic testing of SU in chat-oriented models, through assessment of models' ability to track and enumerate environment states.Our environment also allows for close analysis of dynamics of model performance, to better understand underlying causes for performance patterns.We apply our test to ChatGPT, the stateof-the-art chatbot, and find that despite the fundamental simplicity of the task, the model's performance reflects an inability to retain correct environment states across time.Our follow-up analyses suggest that performance degradation is largely because ChatGPT has non-persistent in-context memory (although it can access the full dialogue history) and it is susceptible to hallucinated updates-including updates that artificially inflate accuracies.Our findings suggest overall that ChatGPT is not currently equipped for robust tracking of situation states, and that trust in the impressive dialogue performance of ChatGPT comes with risks.We release the codebase for reproducing our test environment, as well as all prompts and API responses from ChatGPT, at https://github.com/ yangalan123/SituationalTesting.
Chenghao Yang 0001, Allyson Ettinger
EMNLP2
2023 Linguistic Productivity: the Case of Determiners in English
abstract
Raquel G. Alhama, Ruthe Foushee, Daniel Byrne, Allyson Ettinger, Susan Goldin-Meadow, Afra Alishahi. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Raquel G. Alhama, Ruthe Foushee, Daniel Byrne, Allyson Ettinger, Susan Goldin-Meadow, Afra Alishahi
IJCNLP (1)4
2023 Faith and Fate: Limits of Transformers on Compositionality
abstract
Transformer large language models (LLMs) have sparked admiration for their exceptional performance on tasks that demand intricate multi-step reasoning. Yet, these models simultaneously show failures on surprisingly trivial problems. This begs the question: Are these errors incidental, or do they signal more substantial limitations? In an attempt to demystify transformer LLMs, we investigate the limits of these models across three representative compositional tasks---multi-digit multiplication, logic grid puzzles, and a classic dynamic programming problem. These tasks require breaking problems down into sub-steps and synthesizing these steps into a precise answer. We formulate compositional tasks as computation graphs to systematically quantify the level of complexity, and break down reasoning steps into intermediate sub-procedures. Our empirical findings suggest that transformer LLMs solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem-solving skills. To round off our empirical study, we provide theoretical arguments on abstract multi-step reasoning problems that highlight how autoregressive generations' performance can rapidly decay with increased task complexity.
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Li 0069, Bill Y. Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras 0001, Jena D. Hwang, Soumya Sanyal 0001, Xiang Ren 0001, Allyson Ettinger, Zaïd Harchaoui, Yejin Choi 0001
NeurIPS14
2022 A Property Induction Framework for Neural Language Models
Kanishka Misra, Julia Taylor Rayz, Allyson Ettinger
CogSci3
2022 "No, They Did Not": Dialogue Response Dynamics in Pre-trained Language Models
abstract
A critical component of competence in language is being able to identify relevant components of an utterance and reply appropriately. In this paper we examine the extent of such dialogue response sensitivity in pre-trained language models, conducting a series of experiments with a particular focus on sensitivity to dynamics involving phenomena of at-issueness and ellipsis. We find that models show clear sensitivity to a distinctive role of embedded clauses, and a general preference for responses that target main clause content of prior utterances. However, the results indicate mixed and generally weak trends with respect to capturing the full range of dynamics involved in targeting at-issue versus not-at-issue content. Additionally, models show fundamental limitations in grasp of the dynamics governing ellipsis, and response selections show clear interference from superficial factors that outweigh the influence of principled discourse constraints.
Sanghee J. Kim, Lang Yu, Allyson Ettinger
COLING3
2021 Do language models learn typicality judgments from text?
Kanishka Misra, Allyson Ettinger, Julia Taylor Rayz
CogSci2
2021 Pragmatic competence of pre-trained language models through the lens of discourse connectives
abstract
As pre-trained language models (LMs) continue to dominate NLP, it is increasingly important that we understand the depth of language capabilities in these models.In this paper, we target pre-trained LMs' competence in pragmatics, with a focus on pragmatics relating to discourse connectives.We formulate cloze-style tests using a combination of naturally-occurring data and controlled inputs drawn from psycholinguistics.We focus on testing models' ability to use pragmatic cues to predict discourse connectives, models' ability to understand implicatures relating to connectives, and the extent to which models show humanlike preferences regarding temporal dynamics of connectives.We find that although models predict connectives reasonably well in the context of naturally-occurring data, when we control contexts to isolate high-level pragmatic cues, model sensitivity is much lower.Models also do not show substantial humanlike temporal preferences.Overall, the findings suggest that at present, dominant pretraining paradigms do not result in substantial pragmatic competence in our models.
Lalchand Pandia, Yan Cong, Allyson Ettinger
CoNLL3
2021 Sorting through the noise: Testing robustness of information processing in pre-trained language models
abstract
Pre-trained LMs have shown impressive performance on downstream NLP tasks, but we have yet to establish a clear understanding of their sophistication when it comes to processing, retaining, and applying information presented in their input.In this paper we tackle a component of this question by examining robustness of models' ability to deploy relevant context information in the face of distracting content.We present models with cloze tasks requiring use of critical context information, and introduce distracting content to test how robustly the models retain and use that critical information for prediction.We also systematically manipulate the nature of these distractors, to shed light on dynamics of models' use of contextual cues.We find that although models appear in simple contexts to make predictions based on understanding and applying relevant facts from prior context, the presence of distracting but irrelevant content has clear impact in confusing model predictions.In particular, models appear particularly susceptible to factors of semantic similarity and word position.The findings are consistent with the conclusion that LM predictions are driven in large part by superficial contextual cues, rather than by robust representations of context meaning.
Lalchand Pandia, Allyson Ettinger
EMNLP (1)2
2020 Spying on Your Neighbors: Fine-grained Probing of Contextual Embeddings for Information about Surrounding Words
abstract
Although models using contextual word embeddings have achieved state-of-the-art results on a host of NLP tasks, little is known about exactly what information these embeddings encode about the context words that they are understood to reflect.To address this question, we introduce a suite of probing tasks that enable fine-grained testing of contextual embeddings for encoding of information about surrounding words.We apply these tasks to examine the popular BERT, ELMo and GPT contextual encoders, and find that each of our tested information types is indeed encoded as contextual information across tokens, often with nearperfect recoverability-but the encoders vary in which features they distribute to which tokens, how nuanced their distributions are, and how robust the encoding of each feature is to distance.We discuss implications of these results for how different types of models break down and prioritize word-level context information when constructing token embeddings.
Josef Klafka, Allyson Ettinger
ACL2
2020 PeTra: A Sparsely Supervised Memory Model for People Tracking
abstract
We propose PeTra, a memory-augmented neural network designed to track entities in its memory slots.PeTra is trained using sparse annotation from the GAP pronoun resolution dataset and outperforms a prior memory model on the task while using a simpler architecture.We empirically compare key modeling choices, finding that we can simplify several aspects of the design of the memory module while retaining strong performance.To measure the people tracking capability of memory models, we (a) propose a new diagnostic evaluation based on counting the number of unique entities in text, and (b) conduct a small scale human evaluation to compare evidence of people tracking in the memory logs of PeTra relative to a previous approach.PeTra is highly effective in both evaluations, demonstrating its ability to track people in its memory despite being trained with limited annotation.
Shubham Toshniwal, Allyson Ettinger, Kevin Gimpel, Karen Livescu
ACL2
2020 Exploring Lexical Relations in BERT using Semantic Priming
Kanishka Misra, Allyson Ettinger, Julia Taylor Rayz
CogSci2
2020 Learning to Ignore: Long Document Coreference with Bounded Memory Neural Networks
abstract
Long document coreference resolution remains a challenging task due to the large memory and runtime requirements of current models.Recent work doing incremental coreference resolution using just the global representation of entities shows practical benefits but requires keeping all entities in memory, which can be impractical for long documents.We argue that keeping all entities in memory is unnecessary, and we propose a memoryaugmented neural network that tracks only a small bounded number of entities at a time, thus guaranteeing a linear runtime in length of document.We show that (a) the model remains competitive with models with high memory and computational requirements on OntoNotes and LitBank, and (b) the model learns an efficient memory management strategy easily outperforming a rule-based strategy.
Shubham Toshniwal, Sam Wiseman, Allyson Ettinger, Karen Livescu, Kevin Gimpel
EMNLP (1)3
2020 Assessing Phrasal Representation and Composition in Transformers
abstract
Deep transformer models have pushed performance on NLP tasks to new limits, suggesting sophisticated treatment of complex linguistic inputs, such as phrases.However, we have limited understanding of how these models handle representation of phrases, and whether this reflects sophisticated composition of phrase meaning like that done by humans.In this paper, we present systematic analysis of phrasal representations in state-of-the-art pre-trained transformers.We use tests leveraging human judgments of phrase similarity and meaning shift, and compare results before and after control of word overlap, to tease apart lexical effects versus composition effects.We find that phrase representation in these models relies heavily on word content, with little evidence of nuanced composition.We also identify variations in phrase representation quality across models, layers, and representation types, and make corresponding recommendations for usage of representations from these models.
Lang Yu, Allyson Ettinger
EMNLP (1)2
2020 What BERT is Not: Lessons from a New Suite of Psycholinguistic Diagnostics for Language Models
abstract
Pre-training by language modeling has become a popular and successful approach to NLP tasks, but we have yet to understand exactly what linguistic capacities these pre-training processes confer upon models. In this paper we introduce a suite of diagnostics drawn from human language experiments, which allow us to ask targeted questions about information used by language models for generating predictions in context. As a case study, we apply these diagnostics to the popular BERT model, finding that it can generally distinguish good from bad completions involving shared category or role reversal, albeit with less sensitivity than humans, and it robustly retrieves noun hypernyms, but it struggles with challenging inference and role-based event prediction— and, in particular, it shows clear insensitivity to the contextual impacts of negation.
Allyson Ettinger
Trans. Assoc. Comput. Linguistics1
2018 Assessing Composition in Sentence Vector Representations
abstract
An important component of achieving language understanding is mastering the composition of sentence meaning, but an immediate challenge to solving this problem is the opacity of sentence vector representations produced by current neural sentence composition models. We present a method to address this challenge, developing tasks that directly target compositional meaning information in sentence vector representations with a high degree of precision and control. To enable the creation of these controlled tasks, we introduce a specialized sentence generation system that produces large, annotated sentence sets meeting specified syntactic, semantic and lexical constraints. We describe the details of the method and generation system, and then present results of experiments applying our method to probe for compositional information in embeddings from a number of existing sentence composition models. We find that the method is able to extract useful information about the differing capacities of these models, and we discuss the implications of our results with respect to these systems’ capturing of sentence information. We make available for public use the datasets used for these experiments, as well as the generation system.
Allyson Ettinger, Ahmed Elgohary, Colin Phillips, Philip Resnik
COLING1
2016 Modeling N400 amplitude using vector space models of word representation
Allyson Ettinger, Naomi Feldman, Philip Resnik, Colin Phillips
CogSci1
2016 Retrofitting Sense-Specific Word Vectors Using Parallel Text
abstract
Jauhar et al. (2015) recently proposed to learn sense-specific word representations by "retrofitting" standard distributional word representations to an existing ontology.We observe that this approach does not require an ontology, and can be generalized to any graph defining word senses and relations between them.We create such a graph using translations learned from parallel corpora.On a set of lexical semantic tasks, representations learned using parallel text perform roughly as well as those derived from WordNet, and combining the two representation types significantly improves performance.
Allyson Ettinger, Philip Resnik, Marine Carpuat
HLT-NAACL1
2015 Dialogue focus tracking for zero pronoun resolution
abstract
Sudha Rao, Allyson Ettinger, Hal Daumé III, Philip Resnik. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Sudha Rao, Allyson Ettinger, Hal Daumé III, Philip Resnik
HLT-NAACL2