Daniel Jurafsky

dblp:31/985 · also Dan Jurafsky · DBLP profile ↗
← Back
190ranked-venue papers
14as first author
47since 2021 · last 2026
0000-0002-6459-7745ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 168 · 10 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 8 first-author · 7 since 2021Databases, data management, data science and information retrieval · 11 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-author · 2 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
abstract
Recent evaluations show that large language models (LLMs) frequently fail to challenge users' harmful beliefs in domains ranging from medical advice to social reasoning.We present a unifying analysis through the lens of pragmatics: these safety failures can be understood and addressed as LLMs exhibiting excessive accommodation and insufficient epistemic vigilance.We show that the pragmatic factors affecting accommodation and epistemic vigilance in humans (at-issueness, linguistic encoding, and source reliability) influence LLM behaviors in similar ways.We demonstrate how these factors explain performance differences across three safety benchmarks that test models' ability to challenge harmful beliefs, spanning misinformation (Cancer-Myth, SAGE-Eval) and sycophancy (ELEPHANT).This pragmatic lens further motivates prompting interventions, such as adding the phrase "wait a minute", that drastically improve performance on these difficult benchmarks by shifting pragmatic cues.Our results have practical implications for benchmark design and underscore the importance of pragmatics for understanding model behavior and improving performance.
Myra Cheng, Robert D. Hawkins, Daniel Jurafsky
ACL (1)3
2026 🧑‍🍳 Cooking Up Creativity : Enhancing LLM Creativity through Structured Recombination
abstract
Abstract Large Language Models (LLMs) excel at many tasks, yet they struggle to produce truly creative, diverse ideas. In this paper, we introduce a novel approach that enhances LLM creativity. We apply LLMs for translating between natural language and structured representations, and perform the core creative leap via cognitively inspired manipulations on these representations. Our notion of creativity goes beyond superficial token-level variations; rather, we recombine structured representations of existing ideas, enabling our system to effectively explore a more abstract landscape of ideas. We demonstrate our approach in the culinary domain with DishCover, a model that generates creative recipes. Experiments and domain-expert evaluations reveal that our outputs, which are mostly coherent and feasible, significantly surpass GPT-4o in terms of novelty and diversity, thus outperforming it in creative generation. We hope our work inspires further research into structured creativity in AI.
Moran Mizrahi 0001, Chen Shani, Gabriel Stanovsky, Daniel Jurafsky, Dafna Shahaf
Trans. Assoc. Comput. Linguistics4
2025 HumT DumT: Measuring and controlling human-like language in LLMs
abstract
Should LLMs generate language that makes them seem human?Human-like language might improve user experience, but might also lead to deception, overreliance, and stereotyping.Assessing these potential impacts requires a systematic way to measure human-like tone in LLM outputs.We introduce HUMT and SO-CIOT, metrics for human-like tone and other dimensions of social perceptions in text data based on relative probabilities from an LLM.By measuring HUMT across preference and usage datasets, we find that users prefer less human-like outputs from LLMs in many contexts.HUMT also offers insights into the perceptions and impacts of anthropomorphism: human-like LLM outputs are highly correlated with warmth, social closeness, femininity, and low status, which are closely linked to the aforementioned harms.We introduce DUMT, a method using HUMT to systematically control and reduce the degree of human-like tone while preserving model performance.DUMT offers a practical approach for mitigating risks associated with anthropomorphic language generation.
Myra Cheng, Sunny Yu, Daniel Jurafsky
ACL (1)3
2025 Transcribe, Translate, or Transliterate: An Investigation of Intermediate Representations in Spoken Language Models
Tolúlopé Ògúnrèmí, Christopher D. Manning, Daniel Jurafsky, Karen Livescu
ASRU3
2025 Soft production preferences emerge from a bottleneck on memory
Neil Rathi, Richard Futrell, Daniel Jurafsky
CogSci3
2025 In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
abstract
Human listeners readily adjust to unfamiliar speakers and language varieties through exposure, but do these adaptation benefits extend to state-of-the-art spoken language models (SLMs)?We introduce a scalable framework that allows for in-context learning (ICL) in Phi-4 Multimodal (Phi-4-MM) using interleaved task prompts and audio-text pairs, and find that as few as 12 example utterances (∼50 seconds) at inference time reduce word error rates by a relative 19.7% (1.2 pp.) on average across diverse English corpora.These improvements are most pronounced in lowresource varieties, when the context and target speaker match, and when more examples are provided-though scaling our procedure yields diminishing marginal returns to context length.Overall, we find that our novel ICL adaptation scheme (1) reveals a similar performance profile to human listeners, and (2) demonstrates consistent improvements to automatic speech recognition (ASR) robustness across diverse speakers and language backgrounds.While adaptation succeeds broadly, significant gaps remain for certain varieties, revealing where current models still fall short of human flexibility.We release our prompts and code on GitHub 1
Nathan Roll, Calbert Graham, Yuka Tatsumi, Kim Tien Nguyen, Meghan Sumner, Daniel Jurafsky
EMNLP6
2025 Constructing Datasets From Public Police Body Camera Footage
abstract
The enormous potential of body-worn cameras to improve accountability in policing remains largely unrealized due to large volumes of unreviewed footage. Transcription and diarization tools could aid in reviewing footage, but lack of public data hinders their development. We develop a pipeline to construct public datasets, making use of the small number of videos publicly released by police departments, with capacity to update the data as footage gets released or removed. Our pipeline produces two datasets, a large one with transcriptions automatically extracted from department-generated captions, and a smaller test set where we manually validated transcripts and alignment. We benchmark ASR models, including models fine-tuned on our data, on our test set, to show applications of our datasets and continued challenges of this domain. Our work presents a new vision for leveraging public body-worn camera footage—even when it can’t be rereleased—to help address this critical social issue.
Jamie Rosas-Smith, Martijn Bartelds, Ruizhe Huang, L. Paola García-Perera, Karen Livescu, Daniel Jurafsky, Anjalie Field
ICASSP6
2025 h4rm3l: A Language for Composable Jailbreak Attack Synthesis
abstract
Despite their demonstrated valuable capabilities, state-of-the-art (SOTA) widely deployed large language models (LLMs) still have the potential to cause harm to society due to the ineffectiveness of their safety filters, which can be bypassed by prompt transformations called jailbreak attacks. Current approaches to LLM safety assessment, which employ datasets of templated prompts and benchmarking pipelines, fail to cover sufficiently large and diverse sets of jailbreak attacks, leading to the widespread deployment of unsafe LLMs. Recent research showed that novel jailbreak attacks could be derived by composition; however, a formal composable representation for jailbreak attacks, which, among other benefits, could enable the exploration of a large compositional space of jailbreak attacks through program synthesis methods, has not been previously proposed. We introduce h4rm3l, a novel approach that addresses this gap with a human-readable domain-specific language (DSL). Our framework comprises: (1) The h4rm3l DSL, which formally expresses jailbreak attacks as compositions of parameterized string transformation primitives. (2) A synthesizer with bandit algorithms that efficiently generates jailbreak attacks optimized for a target black box LLM. (3) The h4rm3l red-teaming software toolkit that employs the previous two components and an automated harmful LLM behavior classifier that is strongly aligned with human judgment. We demonstrate h4rm3l's efficacy by synthesizing a dataset of 2656 successful novel jailbreak attacks targeting 6 SOTA open-source and proprietary LLMs (GPT-3.5, GPT-4o, Claude-3-Sonnet, Claude-3-Haiku, Llama-3-8B, and Llama-3-70B), and by benchmarking those models against a subset of these synthesized attacks. Our results show that h4rm3l's synthesized attacks are diverse and more successful than existing jailbreak attacks in literature, with success rates exceeding 90% on SOTA LLMs. Warning: This paper and related research artifacts contain offensive and potentially disturbing prompts and model-generated content.
Moussa Doumbouya, Ananjan Nandi, Gabriel Poesia, Davide Ghilardi, Anna Goldie, Federico Bianchi 0001, Daniel Jurafsky, Christopher D. Manning
ICLR7
2025 What can large language models do for sustainable food?
abstract
Food systems are responsible for a third of human-caused greenhouse gas emissions. We investigate what Large Language Models (LLMs) can contribute to reducing the environmental impacts of food production. We define a typology of design and prediction tasks based on the sustainable food literature and collaboration with domain experts, and evaluate six LLMs on four tasks in our typology. For example, for a sustainable protein design task, food science experts estimated that collaboration with an LLM can reduce time spent by 45% on average, compared to 22% for collaboration with another expert human food scientist. However, for a sustainable menu design task, LLMs produce suboptimal solutions when instructed to consider both human satisfaction and climate impacts. We propose a general framework for integrating LLMs with combinatorial optimization to improve reasoning capabilities. Our approach decreases emissions of food choices by 79% in a hypothetical restaurant while maintaining participants’ satisfaction with their set of choices. Our results demonstrate LLMs’ potential, supported by optimization techniques, to accelerate sustainable food development and adoption.
Anna T. Thomas, Adam Yee, Andrew Mayne, Maya B. Mathur, Daniel Jurafsky, Kristina Gligoric
ICML5
2025 AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders
abstract
Fine-grained steering of language model outputs is essential for safety and reliability. Prompting and finetuning are widely used to achieve these goals, but interpretability researchers have proposed a variety of representation-based techniques as well, including sparse autoencoders (SAEs), linear artificial tomography, supervised steering vectors, linear probes, and representation finetuning. At present, there is no benchmark for making direct comparisons between these proposals. Therefore, we introduce AxBench, a large-scale benchmark for steering and concept detection, and report experiments on Gemma-2-2B and 9B. For steering, we find that prompting outperforms all existing methods, followed by finetuning. For concept detection, representation-based methods such as difference-in-means, perform the best. On both evaluations, SAEs are not competitive. We introduce a novel weakly-supervised representational method (Rank-1 Representation Finetuning; ReFT-r1), which is competitive on both tasks while providing the interpretability advantages that prompting lacks. Along with AxBench, we train and publicly release SAE-scale feature dictionaries for ReFT-r1 and DiffMean.
Zhengxuan Wu, Aryaman Arora, Atticus Geiger, Zheng Wang 0078, Jing Huang 0014, Daniel Jurafsky, Christopher D. Manning, Christopher Potts
ICML6
2025 The ML-SUPERB 2.0 Challenge: Towards Inclusive ASR Benchmarking for All Language Varieties
Chutong Meng, Jiatong Shi, Martijn Bartelds, Shih-Heng Wang, Hsiu-Hsuan Wang, Rafael Mosquera, Sara Hincapie, Daniel Jurafsky, Antonios Anastasopoulos, Hung-yi Lee, Karen Livescu, Shinji Watanabe 0001
INTERSPEECH9
2025 Can Unconfident LLM Annotations Be Used for Confident Conclusions?
abstract
Kristina Gligoric, Tijana Zrnic, Cinoo Lee, Emmanuel Candes, Dan Jurafsky. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kristina Gligoric, Tijana Zrnic, Cinoo Lee, Emmanuel J. Candès, Daniel Jurafsky
NAACL (Long Papers)5
2025 Rethinking Word Similarity: Semantic Similarity through Classification Confusion
abstract
Kaitlyn Zhou, Haishan Gao, Sarah Li Chen, Dan Edelstein, Dan Jurafsky, Chen Shani. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kaitlyn Zhou, Haishan Gao, Sarah Li Chen, Dan Edelstein, Daniel Jurafsky, Chen Shani
NAACL (Long Papers)5
2025 REL-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
abstract
Kaitlyn Zhou, Jena D. Hwang, Xiang Ren, Nouha Dziri, Dan Jurafsky, Maarten Sap. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kaitlyn Zhou, Jena D. Hwang, Xiang Ren 0001, Nouha Dziri, Daniel Jurafsky, Maarten Sap
NAACL (Long Papers)5
2024 CausalGym: Benchmarking causal interpretability methods on linguistic tasks
abstract
Language models (LMs) have proven to be powerful tools for psycholinguistic research, but most prior work has focused on purely behavioural measures (e.g., surprisal comparisons).At the same time, research in model interpretability has begun to illuminate the abstract causal mechanisms shaping LM behavior.To help bring these strands of research closer together, we introduce CausalGym.We adapt and expand the Syntax-Gym suite of tasks to benchmark the ability of interpretability methods to causally affect model behaviour.To illustrate how CausalGym can be used, we study the pythia models (14M-6.9B)and assess the causal efficacy of a wide range of interpretability methods, including linear probing and distributed alignment search (DAS).We find that DAS outperforms the other methods, and so we use it to study the learning trajectory of two difficult linguistic phenomena in pythia-1b: negative polarity item licensing and filler-gap dependencies.Our analysis shows that the mechanism implementing both of these tasks is learned in discrete stages, not gradually.https://github.com/aryamanarora/ causalgym
Aryaman Arora, Daniel Jurafsky, Christopher Potts
ACL (1)2
2024 AnthroScore: A Computational Linguistic Measure of Anthropomorphism
abstract
Myra Cheng, Kristina Gligoric, Tiziano Piccardi, Dan Jurafsky. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Myra Cheng, Kristina Gligoric, Tiziano Piccardi, Daniel Jurafsky
EACL (1)4
2024 Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
abstract
Training large language models to follow instructions makes them perform better on a wide range of tasks and generally become more helpful. However, a perfectly helpful model will follow even the most malicious instructions and readily generate harmful content. In this paper, we raise concerns over the safety of models that only emphasize helpfulness, not harmlessness, in their instruction-tuning. We show that several popular instruction-tuned models are highly unsafe. Moreover, we show that adding just 3\% safety examples (a few hundred demonstrations) when fine-tuning a model like LLaMA can substantially improve its safety. Our safety-tuning does not make models significantly less capable or helpful as measured by standard benchmarks. However, we do find exaggerated safety behaviours, where too much safety-tuning makes models refuse perfectly safe prompts if they superficially resemble unsafe ones. As a whole, our results illustrate trade-offs in training LLMs to be helpful and training them to be safe.
Federico Bianchi 0001, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger, Daniel Jurafsky, Tatsunori B. Hashimoto, James Zou 0001
ICLR5
2024 A Benchmark for Learning to Translate a New Language from One Grammar Book
abstract
Large language models (LLMs) can perform impressive feats with in-context learning or lightweight finetuning. It is natural to wonder how well these models adapt to genuinely new tasks, but how does one find tasks that are unseen in internet-scale training sets? We turn to a field that is explicitly motivated and bottlenecked by a scarcity of web data: low-resource languages. In this paper, we introduce MTOB (Machine Translation from One Book), a benchmark for learning to translate between English and Kalamang—a language with less than 200 speakers and therefore virtually no presence on the web—using several hundred pages of field linguistics reference materials. This task framing is novel in that it asks a model to learn a language from a single human-readable book of grammar explanations, rather than a large mined corpus of in-domain data, more akin to L2 language learning than L1 language acquisition. We demonstrate that baselines using current LLMs are promising but fall short of human performance, achieving 44.7 chrF on Kalamang to English translation and 45.8 chrF on English to Kalamang translation, compared to 51.6 and 57.0 chrF by a human who learned Kalamang from the same reference materials. We hope that MTOB will help measure LLM capabilities along a new dimension, and that the methods developed to solve it could help expand access to language technology for underserved communities by leveraging qualitatively different kinds of data than traditional machine translation.
Garrett Tanzer, Mirac Suzgun, Eline Visser, Daniel Jurafsky, Luke Melas-Kyriazi
ICLR4
2024 How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis
abstract
Negotiation is the basis of social interactions; humans negotiate everything from the price of cars to how to share common resources. With rapidly growing interest in using large language models (LLMs) to act as agents on behalf of human users, such LLM agents would also need to be able to negotiate. In this paper, we study how well LLMs can negotiate with each other. We develop NegotiationArena: a flexible framework for evaluating and probing the negotiation abilities of LLM agents. We implemented three types of scenarios in NegotiationArena to assess LLM's behaviors in allocating shared resources (ultimatum games), aggregate resources (trading games) and buy/sell goods (price negotiations). Each scenario allows for multiple turns of flexible dialogues between LLM agents to allow for more complex negotiations. Interestingly, LLM agents can significantly boost their negotiation outcomes by employing certain behavioral tactics. For example, by pretending to be desolate and desperate, LLMs can improve their payoffs by 20% when negotiating against the standard GPT-4. We also quantify irrational negotiation behaviors exhibited by the LLM agents, many of which also appear in humans. Together, NegotiationArena offers a new environment to investigate LLM interactions, enabling new insights into LLM's theory of mind, irrationality, and reasoning abilities
Federico Bianchi 0001, Patrick John Chia, Mert Yüksekgönül, Jacopo Tagliabue, Daniel Jurafsky, James Zou 0001
ICML5
2024 Model Alignment as Prospect Theoretic Optimization
abstract
Kahneman & Tversky’s $\textit{prospect theory}$ tells us that humans perceive random variables in a biased but well-defined manner (1992); for example, humans are famously loss-averse. We show that objectives for aligning LLMs with human feedback implicitly incorporate many of these biases—the success of these objectives (e.g., DPO) over cross-entropy minimization can partly be ascribed to them belonging to a family of loss functions that we call $\textit{human-aware losses}$ (HALOs). However, the utility functions these methods attribute to humans still differ from those in the prospect theory literature. Using a Kahneman-Tversky model of human utility, we propose a HALO that directly maximizes the utility of generations instead of maximizing the log-likelihood of preferences, as current methods do. We call this approach KTO, and it matches or exceeds the performance of preference-based methods at scales from 1B to 30B, despite only learning from a binary signal of whether an output is desirable. More broadly, our work suggests that there is no one HALO that is universally superior; the best loss depends on the inductive biases most appropriate for a given setting, an oft-overlooked consideration.
Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Daniel Jurafsky, Douwe Kiela
ICML4
2024 Othering and Low Status Framing of Immigrant Cuisines in US Restaurant Reviews and Large Language Models
abstract
Identifying implicit attitudes toward food can help mitigate social prejudice due to food’s pervasive role as a marker of ethnic identity. Stereotypes about food are representational harms that may contribute to racialized discourse and negatively impact economic outcomes for restaurants. Understanding the presence of representational harms in online corpora in particular is important, given the increasing use of large language models (LLMs) for text generation and their tendency to reproduce attitudes in their training data. Through careful linguistic analyses, we evaluate social theories about attitudes toward immigrant cuisine in a large-scale study of framing differences in 2.1M English language Yelp reviews. Controlling for factors such as restaurant price and neighborhood racial diversity, we find that immigrant cuisines are more likely to be othered using socially constructed frames of authenticity (e.g., authentic, traditional), and that non-European cuisines (e.g., Indian, Mexican) in particular are described as more exotic compared to European ones (e.g., French). We also find that non-European cuisines are more likely to be described as cheap and dirty, even after controlling for price, and even among the most expensive restaurants. Finally, we show that reviews generated by LLMs reproduce similar framing tendencies, pointing to the downstream retention of these representational harms. Our results corroborate social theories of gastronomic stereotyping, revealing racialized evaluative processes and linguistic strategies through which they manifest.
Yiwei Luo, Kristina Gligoric, Daniel Jurafsky
ICWSM3
2024 A layer-wise analysis of Mandarin and English suprasegmentals in SSL speech models
Antón de la Fuente, Daniel Jurafsky
INTERSPEECH2
2024 ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
Jiatong Shi, Shih-Heng Wang, Martijn Bartelds, Vanya Bannihatti Kumar, Jinchuan Tian, Xuankai Chang, Daniel Jurafsky, Karen Livescu, Hung-yi Lee, Shinji Watanabe 0001
INTERSPEECH8
2024 NLP Systems That Can't Tell Use from Mention Censor Counterspeech, but Teaching the Distinction Helps
abstract
Kristina Gligoric, Myra Cheng, Lucia Zheng, Esin Durmus, Dan Jurafsky. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Kristina Gligoric, Myra Cheng, Lucia Zheng, Esin Durmus, Daniel Jurafsky
NAACL-HLT5
2024 Grounding Gaps in Language Model Generations
abstract
Omar Shaikh, Kristina Gligoric, Ashna Khetan, Matthias Gerstgrasser, Diyi Yang, Dan Jurafsky. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Omar Shaikh, Kristina Gligoric, Ashna Khetan, Matthias Gerstgrasser, Diyi Yang, Daniel Jurafsky
NAACL-HLT6
2024 ReFT: Representation Finetuning for Language Models
abstract
Parameter-efficient finetuning (PEFT) methods seek to adapt large neural models via updates to a small number of *weights*. However, much prior interpretability work has shown that *representations* encode rich semantic information, suggesting that editing representations might be a more powerful alternative. We pursue this hypothesis by developing a family of **Representation Finetuning (ReFT)** methods. ReFT methods operate on a frozen base model and learn task-specific interventions on hidden representations. We define a strong instance of the ReFT family, Low-rank Linear Subspace ReFT (LoReFT), and we identify an ablation of this method that trades some performance for increased efficiency. Both are drop-in replacements for existing PEFTs and learn interventions that are 15x--65x more parameter-efficient than LoRA. We showcase LoReFT on eight commonsense reasoning tasks, four arithmetic reasoning tasks, instruction-tuning, and GLUE. In all these evaluations, our ReFTs deliver the best balance of efficiency and performance, and almost always outperform state-of-the-art PEFTs. Upon publication, we will publicly release our generic ReFT training library.
Zhengxuan Wu, Aryaman Arora, Zheng Wang 0078, Atticus Geiger, Daniel Jurafsky, Christopher D. Manning, Christopher Potts
NeurIPS5
2023 Making More of Little Data: Improving Low-Resource Automatic Speech Recognition Using Data Augmentation
abstract
The performance of automatic speech recognition (ASR) systems has advanced substantially in recent years, particularly for languages for which a large amount of transcribed speech is available.Unfortunately, for low-resource languages, such as minority languages, regional languages or dialects, ASR performance generally remains much lower.In this study, we investigate whether data augmentation techniques could help improve low-resource ASR performance, focusing on four typologically diverse minority languages or language variants (West Germanic: Gronings, West-Frisian; Malayo-Polynesian: Besemah, Nasal).For all four languages, we examine the use of selftraining, where an ASR system trained with the available human-transcribed data is used to generate transcriptions, which are then combined with the original data to train a new ASR system.For Gronings, for which there was a preexisting text-to-speech (TTS) system available, we also examined the use of TTS to generate ASR training data from text-only sources.We find that using a self-training approach consistently yields improved performance (a relative WER reduction up to 20.5% compared to using an ASR system trained on 24 minutes of manually transcribed speech).The performance gain from TTS augmentation for Gronings was even stronger (up to 25.5% relative reduction in WER compared to a system based on 24 minutes of manually transcribed speech).In sum, our results show the benefit of using selftraining or (if possible) TTS-generated data as an efficient solution to overcome the limitations of data availability for resource-scarce languages in order to improve ASR performance.
Martijn Bartelds, Nay San, Bradley McDonnell, Daniel Jurafsky, Martijn Wieling 0001
ACL (1)4
2023 Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models
abstract
To recognize and mitigate harms from large language models (LLMs), we need to understand the prevalence and nuances of stereotypes in LLM outputs.Toward this end, we present Marked Personas, a prompt-based method to measure stereotypes in LLMs for intersectional demographic groups without any lexicon or data labeling.Grounded in the sociolinguistic concept of markedness (which characterizes explicitly linguistically marked categories versus unmarked defaults), our proposed method is twofold: 1) prompting an LLM to generate personas, i.e., natural language descriptions, of the target demographic group alongside personas of unmarked, default groups; 2) identifying the words that significantly distinguish personas of the target group from corresponding unmarked ones.We find that the portrayals generated by GPT-3.5 and GPT-4 contain higher rates of racial stereotypes than human-written portrayals using the same prompts.The words distinguishing personas of marked (non-white, non-male) groups reflect patterns of othering and exoticizing these demographics.An intersectional lens further reveals tropes that dominate portrayals of marginalized groups, such as tropicalism and the hypersexualization of minoritized women.These representational harms have concerning implications for downstream applications like story generation.
Myra Cheng, Esin Durmus, Daniel Jurafsky
ACL (1)3
2023 Self-Destructing Models: Increasing the Costs of Harmful Dual Uses of Foundation Models
abstract
A growing ecosystem of large, open-source foundation models has reduced the labeled data and technical expertise necessary to apply machine learning to many new problems. Yet foundation models pose a clear dual-use risk, indiscriminately reducing the costs of building both harmful and beneficial machine learning systems. Policy tools such as restricted model access and export controls are the primary methods currently used to mitigate such dual-use risks. In this work, we review potential safe-release strategies and argue that both policymakers and AI researchers would benefit from fundamentally new technologies enabling more precise control over the downstream usage of open-source foundation models. We propose one such approach: the task blocking paradigm, in which foundation models are trained with an additional mechanism to impede adaptation to harmful tasks without sacrificing performance on desirable tasks. We call the resulting models self-destructing models, inspired by mechanisms that prevent adversaries from using tools for harmful purposes. We present an algorithm for training self-destructing models leveraging techniques from meta-learning and adversarial learning, which we call meta-learned adversarial censoring (MLAC). In a small-scale experiment, we show MLAC can largely prevent a BERT-style model from being re-purposed to perform gender identification without harming the model’s ability to perform profession classification.
Peter Henderson 0002, Eric Mitchell, Christopher D. Manning, Daniel Jurafsky, Chelsea Finn
AIES4
2023 When Do Pre-Training Biases Propagate to Downstream Tasks? A Case Study in Text Summarization
abstract
Faisal Ladhak, Esin Durmus, Mirac Suzgun, Tianyi Zhang, Dan Jurafsky, Kathleen McKeown, Tatsunori Hashimoto. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Faisal Ladhak, Esin Durmus, Mirac Suzgun, Daniel Jurafsky, Kathy McKeown, Tatsunori B. Hashimoto
EACL5
2023 Navigating the Grey Area: How Expressions of Uncertainty and Overconfidence Affect Language Models
abstract
The increased deployment of LMs for realworld tasks involving knowledge and facts makes it important to understand model epistemology: what LMs think they know, and how their attitudes toward that knowledge are affected by language use in their inputs.Here, we study an aspect of model epistemology: how epistemic markers of certainty, uncertainty, or evidentiality like "I'm sure it's", "I think it's", or "Wikipedia says it's" affect models, and whether they contribute to model failures.We develop a typology of epistemic markers and inject 50 markers into prompts for question answering.We find that LMs are highly sensitive to epistemic markers in prompts, with accuracies varying more than 80%.Surprisingly, we find that expressions of high certainty result in a 7% decrease in accuracy as compared to low certainty expressions; similarly, factive verbs hurt performance, while evidentials benefit performance.Our analysis of a popular pretraining dataset shows that these markers of uncertainty are associated with answers on question-answering websites, while markers of certainty are associated with questions.These associations may suggest that the behavior of LMs is based on mimicking observed language use, rather than truly reflecting epistemic uncertainty.
Kaitlyn Zhou, Daniel Jurafsky, Tatsunori B. Hashimoto
EMNLP2
2023 When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It?
Mert Yüksekgönül, Federico Bianchi 0001, Pratyusha Kalluri, Daniel Jurafsky, James Zou 0001
ICLR4
2023 Developing Speech Processing Pipelines for Police Accountability
Anjalie Field, Nay San, Jennifer L. Eberhardt, Daniel Jurafsky
INTERSPEECH5
2023 Ecosystem-level Analysis of Deployed Machine Learning Reveals Homogeneous Outcomes
abstract
Machine learning is traditionally studied at the model level: researchers measure and improve the accuracy, robustness, bias, efficiency, and other dimensions of specific models. In practice, however, the societal impact of any machine learning model is partially determined by the context into which it is deployed. To capture this, we introduce *ecosystem-level analysis:* rather than analyzing a single model, we consider the collection of models that are deployed in a given context. For example, ecosystem-level analysis in hiring recognizes that a job candidate’s outcomes are determined not only by a single hiring algorithm or firm but instead by the collective decisions of all the firms to which the candidate applied. Across three modalities (text, images, speech) and 11 datasets, we establish a clear trend: deployed machine learning is prone to systemic failure, meaning some users are exclusively misclassified by all models available. Even when individual models improve at the population level over time, we find these improvements rarely reduce the prevalence of systemic failure. Instead, the benefits of these improvements predominantly accrue to individuals who are already correctly classified by other models. In light of these trends, we analyze medical imaging for dermatology, a setting where the costs of systemic failure are especially high. While traditional analyses reveal that both models and humans exhibit racial performance disparities, ecosystem-level analysis reveals new forms of racial disparity in model predictions that do not present in human predictions. These examples demonstrate that ecosystem-level analysis has unique strengths in characterizing the societal impact of machine learning.
Connor Toups, Rishi Bommasani, Kathleen Creel, Sarah H. Bana, Daniel Jurafsky, Percy Liang
NeurIPS5
2023 Foundation Models and Fair Use
abstract
Existing foundation models are trained on copyrighted material. Deploying these models can pose both legal and ethical risks when data creators fail to receive appropriate attribution or compensation. In the United States and several other countries, copyrighted content may be used to build foundation models without incurring liability due to the fair use doctrine. However, there is a caveat: If the model produces output that is similar to copyrighted data, particularly in scenarios that affect the market of that data, fair use may no longer apply to the output of the model. In this work, we emphasize that fair use is not guaranteed, and additional work may be necessary to keep model development and deployment squarely in the realm of fair use. First, we survey the potential risks of developing and deploying foundation models based on copyrighted content. We review relevant U.S. case law, drawing parallels to existing and potential applications for generating text, source code, and visual art. Experiments confirm that popular foundation models can generate content considerably similar to copyrighted material. Second, we discuss technical mitigations that can help foundation models stay in line with fair use. We argue that more research is needed to align mitigation strategies with the current state of the law. Third, we suggest that the law and technical mitigations should co-evolve. For example, coupled with other policy mechanisms, the law could more explicitly consider safe harbors when strong technical tools are used to mitigate infringement harms. This co-evolution may help strike a balance between intellectual property and innovation, which speaks to the original goal of fair use. But we emphasize that the strategies we describe here are not a panacea and more work is needed to develop policies that address the potential harms of foundation models.
Peter Henderson 0002, Xuechen Li 0005, Daniel Jurafsky, Tatsunori B. Hashimoto, Mark A. Lemley, Percy Liang
J. Mach. Learn. Res.3
2022 The Authenticity Gap in Human Evaluation
abstract
Human ratings are the gold standard in NLG evaluation.The standard protocol is to collect ratings of generated text, average across annotators, and rank NLG systems by their average scores.However, little consideration has been given as to whether this approach faithfully captures human preferences.Analyzing this standard protocol through the lens of utility theory in economics, we identify the implicit assumptions it makes about annotators.These assumptions are often violated in practice, in which case annotator ratings cease to reflect their preferences.The most egregious violations come from using Likert scales, which provably reverse the direction of the true preference in certain cases.We suggest improvements to the standard protocol to make it more theoretically sound, but even in its improved form, it cannot be used to evaluate open-ended tasks like story generation.For the latter, we propose a new human evaluation protocol called systemlevel probabilistic assessment (SPA).When human evaluation of stories is done with SPA, we can recover the ordering of GPT-3 models by size, with statistically significant results.However, when human evaluation is done with the standard protocol, less than half of the expected preferences can be recovered (e.g., there is no significant difference between curie and davinci, despite using a highly powered test).
Kawin Ethayarajh, Daniel Jurafsky
EMNLP2
2022 Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language Models
abstract
We propose a method for arbitrary textual style transfer (TST)-the task of transforming a text into any given style-utilizing general-purpose pre-trained language models.Our method, Prompt-and-Rerank, is based on a mathematical formulation of the TST task, decomposing it into three constituent components: textual similarity, target style strength, and fluency.Our method uses zero-shot or few-shot prompting to obtain a set of candidate generations in the target style, and then re-ranks them according to the three components.Our method enables small pre-trained language models to perform on par with state-of-the-art large-scale models while using two orders of magnitude less compute and memory.We also investigate the effect of model size and prompt design (e.g., prompt paraphrasing and delimiter-pair choice) on style transfer quality across seven diverse textual style transfer datasets, finding, among other things, that delimiter-pair choice has a large impact on performance, and that models have biases on the direction of style transfer.1
Mirac Suzgun, Luke Melas-Kyriazi, Daniel Jurafsky
EMNLP3
2022 Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset
abstract
One concern with the rise of large language models lies with their potential for significant harm, particularly from pretraining on biased, obscene, copyrighted, and private information. Emerging ethical approaches have attempted to filter pretraining material, but such approaches have been ad hoc and failed to take context into account. We offer an approach to filtering grounded in law, which has directly addressed the tradeoffs in filtering material. First, we gather and make available the Pile of Law, a ~256GB (and growing) dataset of open-source English-language legal and administrative data, covering court opinions, contracts, administrative rules, and legislative records. Pretraining on the Pile of Law may help with legal tasks that have the promise to improve access to justice. Second, we distill the legal norms that governments have developed to constrain the inclusion of toxic or private content into actionable lessons for researchers and discuss how our dataset reflects these norms. Third, we show how the Pile of Law offers researchers the opportunity to learn such filtering rules directly from the data, providing an exciting new research direction in model-based processing.
Peter Henderson 0002, Mark S. Krass, Lucia Zheng, Neel Guha, Christopher D. Manning, Daniel Jurafsky, Daniel E. Ho
NeurIPS6
2022 Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?
abstract
As the scope of machine learning broadens, we observe a recurring theme of algorithmic monoculture: the same systems, or systems that share components (e.g. datasets, models), are deployed by multiple decision-makers. While sharing offers advantages like amortizing effort, it also has risks. We introduce and formalize one such risk, outcome homogenization: the extent to which particular individuals or groups experience the same outcomes across different deployments. If the same individuals or groups exclusively experience undesirable outcomes, this may institutionalize systemic exclusion and reinscribe social hierarchy. We relate algorithmic monoculture and outcome homogenization by proposing the component sharing hypothesis: if algorithmic systems are increasingly built on the same data or models, then they will increasingly homogenize outcomes. We test this hypothesis on algorithmic fairness benchmarks, demonstrating that increased data-sharing reliably exacerbates homogenization and individual-level effects generally exceed group-level effects. Further, given the current regime in AI of foundation models, i.e. pretrained models that can be adapted to myriad downstream tasks, we test whether model-sharing homogenizes outcomes across tasks. We observe mixed results: we find that for both vision and language settings, the specific methods for adapting a foundation model significantly influence the degree of outcome homogenization. We also identify societal challenges that inhibit the measurement, diagnosis, and rectification of outcome homogenization in deployed machine learning systems.
Rishi Bommasani, Kathleen Creel, Ananya Kumar, Daniel Jurafsky, Percy Liang
NeurIPS4
2021 Measuring Conversational Uptake: A Case Study on Student-Teacher Interactions
abstract
Dorottya Demszky, Jing Liu, Zid Mancenido, Julie Cohen, Heather Hill, Dan Jurafsky, Tatsunori Hashimoto. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Dorottya Demszky, Jing Liu 0064, Zid Mancenido, Julie Cohen, Heather Hill, Daniel Jurafsky, Tatsunori B. Hashimoto
ACL/IJCNLP (1)6
2021 Leveraging Pre-Trained Representations to Improve Access to Untranscribed Speech from Endangered Languages
abstract
Pre-trained speech representations like wav2vec 2.0 are a powerful tool for automatic speech recognition (ASR). Yet many endangered languages lack sufficient data for pre-training such models, or are predominantly oral vernaculars without a standardised writing system, precluding fine-tuning. Query-by-example spoken term detection (QbE-STD) offers an alternative for iteratively indexing untranscribed speech corpora by locating spoken query terms. Using data from 7 Australian Aboriginal languages and a regional variety of Dutch, all of which are endangered or vulnerable, we show that QbE-STD can be improved by leveraging representations developed for ASR (wav2vec 2.0: the English monolingual model and XLSR53 multilingual model). Surprisingly, the English model outperformed the multilingual model on 4 Australian language datasets, raising questions around how to optimally leverage self-supervised speech representations for QbE-STD. Nevertheless, we find that wav2vec 2.0 representations (either English or XLSR53) offer large improvements (56-86% relative) over state-of-the-art approaches on our endangered language datasets.
Nay San, Martijn Bartelds, Mitchell Browne, Lily Clifford, Fiona Gibson, John Mansfield, David Nash 0002, Jane Simpson, Myfany Turpin, Maria Vollmer, Sasha Wilmoth, Daniel Jurafsky
ASRU12
2021 The Emergence of the Shape Bias Results from Communicative Efficiency
abstract
By the age of two, children tend to assume that new word categories are based on objects' shape, rather than their color or texture; this assumption is called the shape bias. They are thought to learn this bias by observing that their caregiver's language is biased towards shape based categories. This presents a chicken and egg problem: if the shape bias must be present in the language in order for children to learn it, how did it arise in language in the first place? In this paper, we propose that communicative efficiency explains both how the shape bias emerged and why it persists across generations. We model this process with neural emergent language agents that learn to communicate about raw pixelated images. First, we show that the shape bias emerges as a result of efficient communication strategies employed by agents. Second, we show that pressure brought on by communicative need is also necessary for it to persist across generations; simply having a shape bias in an agent's input language is insufficient. These results suggest that, over and above the operation of other learning strategies, the shape bias in human learners may emerge and be sustained by communicative pressures.
Eva Portelance, Michael C. Frank, Daniel Jurafsky, Alessandro Sordoni, Romain Laroche
CoNLL3
2021 Focus on what matters: Applying Discourse Coherence Theory to Cross Document Coreference
abstract
Performing event and entity coreference resolution across documents vastly increases the number of candidate mentions, making it intractable to do the full n 2 pairwise comparisons.Existing approaches simplify by considering coreference only within document clusters, but this fails to handle inter-cluster coreference, common in many applications.As a result cross-document coreference algorithms are rarely applied to downstream tasks.We draw on an insight from discourse coherence theory: potential coreferences are constrained by the reader's discourse focus.We model the entities/events in a reader's focus as a neighborhood within a learned latent embedding space which minimizes the distance between mentions and the centroids of their gold coreference clusters.We then use these neighborhoods to sample only hard negatives to train a fine-grained classifier on mention pairs and their local discourse features.Our approach 1 achieves state-of-the-art results for both events and entities on the ECB+, Gun Violence, Football Coreference, and Cross-Domain Cross-Document Coreference corpora.Furthermore, training on multiple corpora improves average performance across all datasets by 17.2 F1 points, leading to a robust coreference resolution model for use in downstream tasks where link distribution is unknown.
William Barr Held, Dan Iter, Daniel Jurafsky
EMNLP (1)3
2021 Nearest Neighbor Machine Translation
Urvashi Khandelwal, Angela Fan, Daniel Jurafsky, Luke Zettlemoyer, Mike Lewis
ICLR3
2021 Improving Factual Completeness and Consistency of Image-to-Text Radiology Report Generation
abstract
Yasuhide Miura, Yuhao Zhang, Emily Tsai, Curtis Langlotz, Dan Jurafsky. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yasuhide Miura, Yuhao Zhang 0004, Emily Bao Tsai, Curt Langlotz, Daniel Jurafsky
NAACL-HLT5
2021 Causal Effects of Linguistic Properties
abstract
Reid Pryzant, Dallas Card, Dan Jurafsky, Victor Veitch, Dhanya Sridhar. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Reid Pryzant, Dallas Card, Daniel Jurafsky, Victor Veitch, Dhanya Sridhar
NAACL-HLT3
2021 Sensitivity as a Complexity Measure for Sequence Classification Tasks
abstract
Abstract We introduce a theoretical framework for understanding and predicting the complexity of sequence classification tasks, using a novel extension of the theory of Boolean function sensitivity. The sensitivity of a function, given a distribution over input sequences, quantifies the number of disjoint subsets of the input sequence that can each be individually changed to change the output. We argue that standard sequence classification methods are biased towards learning low-sensitivity functions, so that tasks requiring high sensitivity are more difficult. To that end, we show analytically that simple lexical classifiers can only express functions of bounded sensitivity, and we show empirically that low-sensitivity functions are easier to learn for LSTMs. We then estimate sensitivity on 15 NLP tasks, finding that sensitivity is higher on challenging tasks collected in GLUE than on simple text classification tasks, and that sensitivity predicts the performance both of simple lexical classifiers and of vanilla BiLSTMs without pretrained contextualized embeddings. Within a task, sensitivity predicts which inputs are hard for such simple models. Our results suggest that the success of massively pretrained contextual representations stems in part because they provide representations from which information can be extracted by low-sensitivity decoders.
Michael Hahn 0001, Daniel Jurafsky, Richard Futrell
Trans. Assoc. Comput. Linguistics2
2020 Automatically Neutralizing Subjective Bias in Text
abstract
Texts like news, encyclopedias, and some social media strive for objectivity. Yet bias in the form of inappropriate subjectivity — introducing attitudes via framing, presupposing truth, and casting doubt — remains ubiquitous. This kind of bias erodes our collective trust and fuels social conflict. To address this issue, we introduce a novel testbed for natural language generation: automatically bringing inappropriately subjective text into a neutral point of view (“neutralizing” biased text). We also offer the first parallel corpus of biased language. The corpus contains 180,000 sentence pairs and originates from Wikipedia edits that removed various framings, presuppositions, and attitudes from biased sentences. Last, we propose two strong encoder-decoder baselines for the task. A straightforward yet opaque concurrent system uses a BERT encoder to identify subjective words as part of the generation process. An interpretable and controllable modular algorithm separates these steps, using (1) a BERT-based classifier to identify problematic words and (2) a novel join embedding through which the classifier can edit the hidden states of the encoder. Large-scale human evaluation across four domains (encyclopedias, news headlines, books, and political speeches) suggests that these algorithms are a first step towards the automatic identification and reduction of bias.
Reid Pryzant, Richard Diehl Martinez, Nathan Dass, Sadao Kurohashi, Daniel Jurafsky, Diyi Yang
AAAI5
2020 Pretraining with Contrastive Sentence Objectives Improves Discourse Performance of Language Models
abstract
Recent models for unsupervised representation learning of text have employed a number of techniques to improve contextual word representations but have put little focus on discourse-level representations.We propose CONPONO 1 , an inter-sentence objective for pretraining language models that models discourse coherence and the distance between sentences.Given an anchor sentence, our model is trained to predict the text k sentences away using a sampled-softmax objective where the candidates consist of neighboring sentences and sentences randomly sampled from the corpus.On the discourse representation benchmark DiscoEval, our model improves over the previous state-of-the-art by up to 13% and on average 4% absolute across 7 tasks.Our model is the same size as BERT-Base, but outperforms the much larger BERT-Large model and other more recent approaches that incorporate discourse.We also show that CONPONO yields gains of 2%-6% absolute even for tasks that do not explicitly evaluate discourse: textual entailment (RTE), common sense reasoning (COPA) and reading comprehension (ReCoRD).
Dan Iter, Kelvin Guu, Larry Lansing, Daniel Jurafsky
ACL4
2020 Social Bias Frames: Reasoning about Social and Power Implications of Language
abstract
contains content that may be offensive or upsetting.
Maarten Sap, Saadia Gabriel, Lianhui Qin, Daniel Jurafsky, Noah A. Smith, Yejin Choi 0001
ACL4
2020 With Little Power Comes Great Responsibility
abstract
Despite its importance to experimental design, statistical power (the probability that, given a real effect, an experiment will reject the null hypothesis) has largely been ignored by the NLP community.Underpowered experiments make it more difficult to discern the difference between statistical noise and meaningful model improvements, and increase the chances of exaggerated findings.By metaanalyzing a set of existing NLP papers and datasets, we characterize typical power for a variety of settings and conclude that underpowered experiments are common in the NLP literature.In particular, for several tasks in the popular GLUE benchmark, small test sets mean that most attempted comparisons to state of the art models will not be adequately powered.Similarly, based on reasonable assumptions, we find that the most typical experimental design for human rating studies will be underpowered to detect small model differences, of the sort that are frequently studied.For machine translation, we find that typical test sets of 2000 sentences have approximately 75% power to detect differences of 1 BLEU point.To improve the situation going forward, we give an overview of best practices for power analysis in NLP and release a series of notebooks to assist with future power analyses.1
Dallas Card, Peter Henderson 0002, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, Daniel Jurafsky
EMNLP (1)6
2020 Utility is in the Eye of the User: A Critique of NLP Leaderboards
abstract
Benchmarks such as GLUE have helped drive advances in NLP by incentivizing the creation of more accurate models.While this leaderboard paradigm has been remarkably successful, a historical focus on performance-based evaluation has been at the expense of other qualities that the NLP community values in models, such as compactness, fairness, and energy efficiency.In this opinion paper, we study the divergence between what is incentivized by leaderboards and what is useful in practice through the lens of microeconomic theory.We frame both the leaderboard and NLP practitioners as consumers and the benefit they get from a model as its utility to them.With this framing, we formalize how leaderboards -in their current form -can be poor proxies for the NLP community at large.For example, a highly inefficient model would provide less utility to practitioners but not to a leaderboard, since it is a cost that only the former must bear.To allow practitioners to better estimate a model's utility to them, we advocate for more transparency on leaderboards, such as the reporting of statistics that are of practical concern (e.g., model size, energy efficiency, and inference latency).
Kawin Ethayarajh, Daniel Jurafsky
EMNLP (1)2
2020 Learning Music Helps You Read: Using Transfer to Study Linguistic Structure in Language Models
abstract
We propose transfer learning as a method for analyzing the encoding of grammatical structure in neural language models.We train LSTMs on non-linguistic data and evaluate their performance on natural language to assess which kinds of data induce generalizable structural features that LSTMs can use for natural language.We find that training on nonlinguistic data with latent structure (MIDI music or Java code) improves test performance on natural language, despite no overlap in surface form or vocabulary.To pinpoint the kinds of abstract structure that models may be encoding to lead to this improvement, we run similar experiments with two artificial parentheses languages: one which has a hierarchical recursive structure, and a control which has paired tokens but no recursion.Surprisingly, training a model on either of these artificial languages leads to the same substantial gains when testing on natural language.Further experiments on transfer between natural languages controlling for vocabulary overlap show that zeroshot performance on a test language is highly correlated with typological syntactic similarity to the training language, suggesting that representations induced by pre-training correspond to the cross-linguistic syntactic properties.Our results provide insights into the ways that neural models represent abstract syntactic structure, and also about the kind of structural inductive biases which allow for natural language acquisition.
Isabel Papadimitriou, Daniel Jurafsky
EMNLP (1)2
2020 Generalization through Memorization: Nearest Neighbor Language Models
Urvashi Khandelwal, Omer Levy, Daniel Jurafsky, Luke Zettlemoyer, Mike Lewis
ICLR3
2020 Language Through a Prism: A Spectral Approach for Multiscale Language Representations
abstract
Language exhibits structure at a wide range of scales, from subwords to words, sentences, paragraphs, and documents. We propose building models that isolate scale-specific information in deep representations, and develop methods for encouraging models during training to learn more about particular scales of interest. Our method for creating scale-specific neurons in deep NLP models constrains how the activation of a neuron can change across the tokens of an input by interpreting those activations as a digital signal and filtering out parts of its frequency spectrum. This technique enables us to extract scale-specific information from BERT representations: by filtering out different frequencies we can produce new representations that perform well on part of speech tagging (word-level), dialog speech acts classification (utterance-level), or topic classification (document-level), while performing poorly on the other tasks. We also present a prism layer for use during training, which constrains different neurons of a BERT model to different parts of the frequency spectrum. Our proposed BERT + Prism model is better able to predict masked tokens using long-range context, and produces individual multiscale representations that perform with comparable or improved performance across all three tasks. Our methods are general and readily applicable to other domains besides language, such as images, audio, and video.
Alex Tamkin, Daniel Jurafsky, Noah D. Goodman
NeurIPS2
2019 Seekers, Providers, Welcomers, and Storytellers: Modeling Social Roles in Online Health Communities
abstract
Participants in online communities often enact different roles when participating in their communities. For example, some in cancer support communities specialize in providing disease-related information or socializing new members. This work clusters the behavioral patterns of users of a cancer support community into specific functional roles. Based on a series of quantitative and qualitative evaluations, this research identified eleven roles that members occupy, such as welcomer and story sharer. We investigated role dynamics, including how roles change over members' lifecycles, and how roles predict long-term participation in the community. We found that members frequently change roles over their history, from ones that seek resources to ones offering help, while the distribution of roles is stable over the community's history. Adopting certain roles early on predicts members' continued participation in the community. Our methodology will be useful for facilitating better use of members' skills and interests in support of community-building efforts.
Diyi Yang, Robert E. Kraut, Tenbroeck Smith, Elijah Mayfield, Daniel Jurafsky
CHI5
2019 Integrating Text and Image: Determining Multimodal Document Intent in Instagram Posts
abstract
Julia Kruk, Jonah Lubin, Karan Sikka, Xiao Lin, Dan Jurafsky, Ajay Divakaran. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Julia Kruk, Jonah Lubin, Karan Sikka, Daniel Jurafsky, Ajay Divakaran
EMNLP/IJCNLP (1)5
2018 Sharp Nearby, Fuzzy Far Away: How Neural Language Models Use Context
abstract
We know very little about how neural language models (LM) use prior linguistic context.In this paper, we investigate the role of context in an LSTM LM, through ablation studies.Specifically, we analyze the increase in perplexity when prior context words are shuffled, replaced, or dropped.On two standard datasets, Penn Treebank and WikiText-2, we find that the model is capable of using about 200 tokens of context on average, but sharply distinguishes nearby context (recent 50 tokens) from the distant history.The model is highly sensitive to the order of words within the most recent sentence, but ignores word order in the long-range context (beyond 50 tokens), suggesting the distant past is modeled only as a rough semantic field or topic.We further find that the neural caching model (Grave et al., 2017b) especially helps the LSTM to copy words from within this distant context.Overall, our analysis not only provides a better understanding of how neural LMs use their context, but also sheds light on recent success from cache-based models.
Urvashi Khandelwal, He He 0001, Peng Qi 0003, Daniel Jurafsky
ACL (1)4
2018 An Information-Theoretic Explanation of Adjective Ordering Preferences
Michael Hahn 0001, Judith Degen, Noah D. Goodman, Daniel Jurafsky, Richard Futrell
CogSci4
2018 Framing and Agenda-Setting in Russian News: a Computational Analysis of Intricate Political Strategies
abstract
Amidst growing concern over media manipulation, NLP attention has focused on overt strategies like censorship and "fake news".Here, we draw on two concepts from the political science literature to explore subtler strategies for government media manipulation: agenda-setting (selecting what topics to cover) and framing (deciding how topics are covered).We analyze 13 years (100K articles) of the Russian newspaper Izvestia and identify a strategy of distraction: articles mention the U.S. more frequently in the month directly following an economic downturn in Russia.We introduce embedding-based methods for cross-lingually projecting English frames to Russian, and discover that these articles emphasize U.S. moral failings and threats to the U.S. Our work offers new ways to identify subtle media manipulation strategies at the intersection of agenda-setting and framing.
Anjalie Field, Doron Kliger, Shuly Wintner, Jennifer Pan, Daniel Jurafsky, Yulia Tsvetkov
EMNLP5
2018 Textual Analogy Parsing: What's Shared and What's Compared among Analogous Facts
abstract
To understand a sentence like "whereas only 10% of White Americans live at or below the poverty line, 28% of African Americans do" it is important not only to identify individual facts, e.g., poverty rates of distinct demographic groups, but also the higher-order relations between them, e.g., the disparity between them.In this paper, we propose the task of Textual Analogy Parsing (TAP) to model this higher-order meaning.The output of TAP is a frame-style meaning representation which explicitly specifies what is shared (e.g., poverty rates) and what is compared (e.g., White Americans vs. African Americans, 10% vs. 28%) between its component facts.Such a meaning representation can enable new applications that rely on discourse understanding such as automated chart generation from quantitative text.We present a new dataset for TAP, baselines, and a model that successfully uses an ILP to enforce the structural constraints of the problem.
Matthew Lamm, Arun Tejasvi Chaganty, Christopher D. Manning, Daniel Jurafsky, Percy Liang
EMNLP4
2018 JESC: Japanese-English Subtitle Corpus
Reid Pryzant, Youngjoo Chung, Daniel Jurafsky, Denny Britz
LREC3
2018 RtGender: A Corpus for Studying Differential Responses to Gender
Rob Voigt, David Jurgens, Vinodkumar Prabhakaran, Daniel Jurafsky, Yulia Tsvetkov
LREC4
2018 Deconfounded Lexicon Induction for Interpretable Social Science
abstract
Reid Pryzant, Kelly Shen, Dan Jurafsky, Stefan Wagner. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Reid Pryzant, Kelly Shen, Daniel Jurafsky, Stefan Wagner 0015
NAACL-HLT3
2018 Noising and Denoising Natural Language: Diverse Backtranslation for Grammar Correction
abstract
Ziang Xie, Guillaume Genthial, Stanley Xie, Andrew Ng, Dan Jurafsky. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Ziang Xie, Guillaume Genthial, Stanley Xie, Andrew Y. Ng, Daniel Jurafsky
NAACL-HLT5
2018 Embedding Logical Queries on Knowledge Graphs
abstract
Learning low-dimensional embeddings of knowledge graphs is a powerful approach used to predict unobserved or missing edges between entities. However, an open challenge in this area is developing techniques that can go beyond simple edge prediction and handle more complex logical queries, which might involve multiple unobserved edges, entities, and variables. For instance, given an incomplete biological knowledge graph, we might want to predict "em what drugs are likely to target proteins involved with both diseases X and Y?" -- a query that requires reasoning about all possible proteins that might interact with diseases X and Y. Here we introduce a framework to efficiently make predictions about conjunctive logical queries -- a flexible but tractable subset of first-order logic -- on incomplete knowledge graphs. In our approach, we embed graph nodes in a low-dimensional space and represent logical operators as learned geometric operations (e.g., translation, rotation) in this embedding space. By performing logical operations within a low-dimensional embedding space, our approach achieves a time complexity that is linear in the number of query variables, compared to the exponential complexity required by a naive enumeration-based approach. We demonstrate the utility of this framework in two application studies on real-world datasets with millions of relations: predicting logical relationships in a network of drug-gene-disease interactions and in a graph-based representation of social interactions derived from a popular web forum.
William L. Hamilton, Payal Bajaj, Marinka Zitnik, Daniel Jurafsky, Jure Leskovec
NeurIPS4
2018 Community Interaction and Conflict on the Web
abstract
Users organize themselves into communities on web platforms. These communities can interact with one another, often leading to conflicts and toxic interactions. However, little is known about the mechanisms of interactions between communities and how they impact users.
Srijan Kumar, William L. Hamilton, Jure Leskovec, Daniel Jurafsky
WWW4
2018 Measuring the Evolution of a Scientific Field through Citation Frames
abstract
Citations have long been used to characterize the state of a scientific field and to identify influential works. However, writers use citations for different purposes, and this varied purpose influences uptake by future scholars. Unfortunately, our understanding of how scholars use and frame citations has been limited to small-scale manual citation analysis of individual papers. We perform the largest behavioral study of citations to date, analyzing how scientific works frame their contributions through different types of citations and how this framing affects the field as a whole. We introduce a new dataset of nearly 2,000 citations annotated for their function, and use it to develop a state-of-the-art classifier and label the papers of an entire field: Natural Language Processing. We then show how differences in framing affect scientific uptake and reveal the evolution of the publication venues and the field as a whole. We demonstrate that authors are sensitive to discourse structure and publication venue when citing, and that how a paper frames its work through citations is predictive of the citation count it will receive. Finally, we use changes in citation framing to show that the field of NLP is undergoing a significant increase in consensus.
David Jurgens, Srijan Kumar, Raine Hoover, Daniel A. McFarland, Daniel Jurafsky
Trans. Assoc. Comput. Linguistics5
2018 Detecting Institutional Dialog Acts in Police Traffic Stops
abstract
We apply computational dialog methods to police body-worn camera footage to model conversations between police officers and community members in traffic stops. Relying on the theory of institutional talk, we develop a labeling scheme for police speech during traffic stops, and a tagger to detect institutional dialog acts (Reasons, Searches, Offering Help) from transcribed text at the turn (78% F-score) and stop (89% F-score) level. We then develop speech recognition and segmentation algorithms to detect these acts at the stop level from raw camera audio (81% F-score, with even higher accuracy for crucial acts like conveying the reason for the stop). We demonstrate that the dialog structures produced by our tagger could reveal whether officers follow law enforcement norms like introducing themselves, explaining the reason for the stop, and asking permission for searches. This work may therefore inform and aid efforts to ensure the procedural justice of police-community interactions.
Vinodkumar Prabhakaran, Camilla Griffiths, Hang Su 0011, Nelson Morgan, Jennifer L. Eberhardt, Daniel Jurafsky
Trans. Assoc. Comput. Linguistics7
2017 A Two-stage Sieve Approach for Quote Attribution
abstract
We present a deterministic sieve-based system for attributing quotations in literary text and a new dataset: QuoteLi3 1 .Quote attribution, determining who said what in a given text, is important for tasks like creating dialogue systems, and in newer areas like computational literary studies, where it creates opportunities to analyze novels at scale rather than only a few at a time.We release QuoteLi3, which contains more than 6,000 annotations linking quotes to speaker mentions and quotes to speaker entities, and introduce a new algorithm for quote attribution.Our twostage algorithm first links quotes to mentions, then mentions to entities.Using two stages encapsulates difficult sub-problems and improves system performance.The modular design allows us to tune either for overall performance or for the high precision appropriate for many use cases.Our system achieves an average F-score of 87.5% across three novels, outperforming previous systems, and can be tuned for precision of 90.4% at a recall of 65.1%.
Grace Muzny, Michael Fang, Angel X. Chang, Daniel Jurafsky
EACL (1)4
2017 Neural Net Models of Open-domain Discourse Coherence
abstract
Discourse coherence is strongly associated with text quality, making it important to natural language generation and understanding.Yet existing models of coherence focus on measuring individual aspects of coherence (lexical overlap, rhetorical structure, entity centering) in narrow domains.In this paper, we describe domainindependent neural models of discourse coherence that are capable of measuring multiple aspects of coherence in existing sentences and can maintain coherence while generating new sentences.We study both discriminative models that learn to distinguish coherent from incoherent discourse, and generative models that produce coherent text, including a novel neural latentvariable Markovian generative model that captures the latent discourse dependencies between sentences in a text.Our work achieves state-of-the-art performance on multiple coherence evaluations, and marks an initial step in generating coherent texts given discourse contexts.
Jiwei Li 0001, Daniel Jurafsky
EMNLP2
2017 Adversarial Learning for Neural Dialogue Generation
abstract
In this paper, drawing intuition from the Turing test, we propose using adversarial training for open-domain dialogue generation: the system is trained to produce sequences that are indistinguishable from human-generated dialogue utterances.We cast the task as a reinforcement learning (RL) problem where we jointly train two systems, a generative model to produce response sequences, and a discriminator-analagous to the human evaluator in the Turing test-to distinguish between the human-generated dialogues and the machine-generated ones.The outputs from the discriminator are then used as rewards for the generative model, pushing the system to generate dialogues that mostly resemble human dialogues.In addition to adversarial training we describe a model for adversarial evaluation that uses success in fooling an adversary as a dialogue evaluation metric, while avoiding a number of potential pitfalls.Experimental results on several metrics, including adversarial evaluation, demonstrate that the adversarially-trained system generates higher-quality responses than previous baselines.
Jiwei Li 0001, Will Monroe, Tianlin Shi, Sébastien Jean, Alan Ritter, Daniel Jurafsky
EMNLP6
2017 Data Noising as Smoothing in Neural Network Language Models
Ziang Xie, Sida I. Wang, Jiwei Li 0001, Daniel Levy 0002, Aiming Nie, Daniel Jurafsky, Andrew Y. Ng
ICLR (Poster)6
2017 Loyalty in Online Communities
William L. Hamilton, Justine Zhang, Cristian Danescu-Niculescu-Mizil, Daniel Jurafsky, Jure Leskovec
ICWSM4
2017 Community Identity and User Engagement in a Multi-Community Landscape
Justine Zhang, William L. Hamilton, Cristian Danescu-Niculescu-Mizil, Daniel Jurafsky, Jure Leskovec
ICWSM4
2017 Building DNN acoustic models for large vocabulary speech recognition
Andrew L. Maas, Peng Qi 0003, Ziang Xie, Awni Y. Hannun, Christopher T. Lengerich, Daniel Jurafsky, Andrew Y. Ng
Comput. Speech Lang.6
2017 A scaffolding approach to coreference resolution integrating statistical and rule-based models
abstract
Abstract We describe a scaffolding approach to the task of coreference resolution that incrementally combines statistical classifiers, each designed for a particular mention type, with rule-based models (for sub-tasks well-matched to determinism). We motivate our design by an oracle-based analysis of errors in a rule-based coreference resolution system, showing that rule-based approaches are poorly suited to tasks that require a large lexical feature space, such as resolving pronominal and common-noun mentions. Our approach combines many advantages: it incrementally builds clusters integrating joint information about entities, uses rules for deterministic phenomena, and integrates rich lexical, syntactic, and semantic features with random forest classifiers well-suited to modeling the complex feature interactions that are known to characterize the coreference task. We demonstrate that all these decisions are important. The resulting system achieves 63.2 F1 on the CoNLL-2012 shared task dataset, outperforming the rule-based starting point by over seven F1 points. Similarly, our system outperforms an equivalent sieve-based approach that relies on logistic regression classifiers instead of random forests by over four F1 points. Lastly, we show that by changing the coreference resolution system from relying on constituent-based syntax to using dependency syntax, which can be generated in linear time, we achieve a runtime speedup of 550 per cent without considerable loss of accuracy.
Heeyoung Lee 0004, Mihai Surdeanu, Daniel Jurafsky
Nat. Lang. Eng.3
2016 Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change
abstract
Understanding how words change their meanings over time is key to models of language and cultural evolution, but historical data on meaning is scarce, making theories hard to develop and test.Word embeddings show promise as a diachronic tool, but have not been carefully evaluated.We develop a robust methodology for quantifying semantic change by evaluating word embeddings (PPMI, SVD, word2vec) against known historical changes.We then use this methodology to reveal statistical laws of semantic evolution.Using six historical corpora spanning four languages and two centuries, we propose two quantitative laws of semantic change: (i) the law of conformity-the rate of semantic change scales with an inverse power-law of word frequency; (ii) the law of innovation-independent of frequency, words that are more polysemous have higher rates of semantic change.
William L. Hamilton, Jure Leskovec, Daniel Jurafsky
ACL (1)3
2016 Predicting the Rise and Fall of Scientific Topics from Trends in their Rhetorical Framing
abstract
Computationally modeling the evolution of science by tracking how scientific topics rise and fall over time has important implications for research funding and public policy.However, little is known about the mechanisms underlying topic growth and decline.We investigate the role of rhetorical framing: whether the rhetorical role or function that authors ascribe to topics (as methods, as goals, as results, etc.) relates to the historical trajectory of the topics.We train topic models and a rhetorical function classifier to map topic models onto their rhetorical roles in 2.4 million abstracts from the Web of Science from 1991-2010.We find that a topic's rhetorical function is highly predictive of its eventual growth or decline.For example, topics that are rhetorically described as results tend to be in decline, while topics that function as methods tend to be in early phases of growth.
Vinodkumar Prabhakaran, William L. Hamilton, Daniel A. McFarland, Daniel Jurafsky
ACL (1)4
2016 Inducing Domain-Specific Sentiment Lexicons from Unlabeled Corpora
abstract
A word's sentiment depends on the domain in which it is used. Computational social science research thus requires sentiment lexicons that are specific to the domains being studied. We combine domain-specific word embeddings with a label propagation framework to induce accurate domain-specific sentiment lexicons using small sets of seed words. We show that our approach achieves state-of-the-art performance on inducing sentiment lexicons from domain-specific corpora and that our purely corpus-based approach outperforms methods that rely on hand-curated resources (e.g., WordNet). Using our framework, we induce and release historical sentiment lexicons for 150 years of English and community-specific sentiment lexicons for 250 online communities from the social media forum Reddit. The historical lexicons we induce show that more than 5% of sentiment-bearing (non-neutral) English words completely switched polarity during the last 150 years, and the community-specific lexicons highlight how sentiment varies drastically between different communities.
William L. Hamilton, Kevin Clark, Jure Leskovec, Daniel Jurafsky
EMNLP4
2016 Cultural Shift or Linguistic Drift? Comparing Two Computational Measures of Semantic Change
abstract
Words shift in meaning for many reasons, including cultural factors like new technologies and regular linguistic processes like subjectification.Understanding the evolution of language and culture requires disentangling these underlying causes.Here we show how two different distributional measures can be used to detect two different types of semantic change.The first measure, which has been used in many previous works, analyzes global shifts in a word's distributional semantics; it is sensitive to changes due to regular processes of linguistic drift, such as the semantic generalization of promise ("I promise."→"Itpromised to be exciting.").The second measure, which we develop here, focuses on local changes to a word's nearest semantic neighbors; it is more sensitive to cultural shifts, such as the change in the meaning of cell ("prison cell" → "cell phone").Comparing measurements made by these two methods allows researchers to determine whether changes are more cultural or linguistic in nature, a distinction that is essential for work in the digital humanities and historical linguistics.
William L. Hamilton, Jure Leskovec, Daniel Jurafsky
EMNLP3
2016 Distinguishing Past, On-going, and Future Events: The EventStatus Corpus
abstract
Determining whether a major societal event has already happened, is still on-going, or may occur in the future is crucial for event prediction, timeline generation, and news summarization.We introduce a new task and a new corpus, EventStatus, which has 4500 English and Spanish articles about civil unrest events labeled as PAST, ON-GOING, or FU-TURE.We show that the temporal status of these events is difficult to classify because local tense and aspect cues are often lacking, time expressions are insufficient, and the linguistic contexts have rich semantic compositionality.We explore two approaches for event status classification: (1) a feature-based SVM classifier augmented with a novel induced lexicon of future-oriented verbs, such as "threatened" and "planned", and (2) a convolutional neural net.Both types of classifiers improve event status recognition over a state-of-the-art TempEval model, and our analysis offers linguistic insights into the semantic compositionality challenges for this new task.
Ruihong Huang, Ignacio Cases, Daniel Jurafsky, Cleo Condoravdi, Ellen Riloff
EMNLP3
2016 Deep Reinforcement Learning for Dialogue Generation
abstract
Recent neural models of dialogue generation offer great promise for generating responses for conversational agents, but tend to be shortsighted, predicting utterances one at a time while ignoring their influence on future outcomes.Modeling the future direction of a dialogue is crucial to generating coherent, interesting dialogues, a need which led traditional NLP models of dialogue to draw on reinforcement learning.In this paper, we show how to integrate these goals, applying deep reinforcement learning to model future reward in chatbot dialogue.The model simulates dialogues between two virtual agents, using policy gradient methods to reward sequences that display three useful conversational properties: informativity, coherence, and ease of answering (related to forward-looking function).We evaluate our model on diversity, length as well as with human judges, showing that the proposed algorithm generates more interactive responses and manages to foster a more sustained conversation in dialogue simulation.This work marks a first step towards learning a neural conversational model based on the long-term success of dialogues.
Jiwei Li 0001, Will Monroe, Alan Ritter, Daniel Jurafsky, Michel Galley, Jianfeng Gao 0001
EMNLP4
2016 Ketchup, Interdisciplinarity, and the Spread of Innovation in Speech and Language Processing
Daniel Jurafsky
INTERSPEECH1
2016 Between- and Within-Speaker Effects of Bilingualism on F0 Variation
Rob Voigt, Daniel Jurafsky, Meghan Sumner
INTERSPEECH2
2016 Visualizing and Understanding Neural Models in NLP
abstract
While neural networks have been successfully applied to many NLP tasks the resulting vectorbased models are very difficult to interpret.For example it's not clear how they achieve compositionality, building sentence meaning from the meanings of words and phrases.In this paper we describe strategies for visualizing compositionality in neural models for NLP, inspired by similar work in computer vision.We first plot unit values to visualize compositionality of negation, intensification, and concessive clauses, allowing us to see wellknown markedness asymmetries in negation.We then introduce methods for visualizing a unit's salience, the amount that it contributes to the final composed meaning from first-order derivatives.Our general-purpose methods may have wide applications for understanding compositionality and other semantic properties of deep networks.
Jiwei Li 0001, Xinlei Chen, Eduard H. Hovy, Daniel Jurafsky
HLT-NAACL4
2015 A Hierarchical Neural Autoencoder for Paragraphs and Documents
abstract
Jiwei Li, Thang Luong, Dan Jurafsky. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Jiwei Li 0001, Minh-Thang Luong, Daniel Jurafsky
ACL (1)3
2015 Do Multi-Sense Embeddings Improve Natural Language Understanding?
abstract
Learning a distinct representation for each sense of an ambiguous word could lead to more powerful and fine-grained mod-els of vector-space representations. Yet while ‘multi-sense ’ methods have been proposed and tested on artificial word-similarity tasks, we don’t know if they im-prove real natural language understanding tasks. In this paper we introduce a multi-sense embedding model based on Chinese Restaurant Processes that achieves state of the art performance on matching human word similarity judgments, and propose a pipelined architecture for incorporating multi-sense embeddings into language un-derstanding. We then test the performance of our model on part-of-speech tagging, named entity recognition, sentiment analysis, semantic relation identification and semantic relat-edness, controlling for embedding dimen-sionality. We find that multi-sense embed-dings do improve performance on some tasks (part-of-speech tagging, semantic re-lation identification, semantic relatedness) but not on others (named entity recogni-tion, various forms of sentiment analysis). We discuss how these differences may be caused by the different role of word sense information in each of the tasks. The re-sults highlight the importance of testing embedding models in real applications. 1
Jiwei Li 0001, Daniel Jurafsky
EMNLP2
2015 When Are Tree Structures Necessary for Deep Learning of Representations?
abstract
Recursive neural models, which use syntactic parse trees to recursively generate representations bottom-up, are a popular architecture.However there have not been rigorous evaluations showing for exactly which tasks this syntax-based method is appropriate.In this paper, we benchmark recursive neural models against sequential recurrent neural models, enforcing applesto-apples comparison as much as possible.We investigate 4 tasks: (1) sentiment classification at the sentence level and phrase level; (2) matching questions to answerphrases; (3) discourse parsing; (4) semantic relation extraction.Our goal is to understand better when, and why, recursive models can outperform simpler models.We find that recursive models help mainly on tasks (like semantic relation extraction) that require longdistance connection modeling, particularly on very long sequences.We then introduce a method for allowing recurrent models to achieve similar performance: breaking long sentences into clause-like units at punctuation and processing them separately before combining.Our results thus help understand the limitations of both classes of models, and suggest directions for improving recurrent models.
Jiwei Li 0001, Thang Luong, Daniel Jurafsky, Eduard H. Hovy
EMNLP3
2015 Lexicon-Free Conversational Speech Recognition with Neural Networks
abstract
We present an approach to speech recognition that uses only a neural network to map acoustic input to characters, a character-level language model, and a beam search decoding procedure.This approach eliminates much of the complex infrastructure of modern speech recognition systems, making it possible to directly train a speech recognizer using errors generated by spoken language understanding tasks.The system naturally handles out of vocabulary words and spoken word fragments.We demonstrate our approach using the challenging Switchboard telephone conversation transcription task, achieving a word error rate competitive with existing baseline systems.To our knowledge, this is the first entirely neural-network-based system to achieve strong speech transcription results on a conversational speech task.We analyze qualitative differences between transcriptions produced by our lexicon-free approach and transcriptions produced by a standard speech recognition system.Finally, we evaluate the impact of large context neural network character language models as compared to standard n-gram models within our framework.
Andrew L. Maas, Ziang Xie, Daniel Jurafsky, Andrew Y. Ng
HLT-NAACL3
2014 Easy does it: more usable CAPTCHAs
abstract
Websites present users with puzzles called CAPTCHAs to curb abuse caused by computer algorithms masquerading as people. While CAPTCHAs are generally effective at stopping abuse, they might impair website usability if they are not properly designed. In this paper we describe how we designed two new CAPTCHA schemes for Google that focus on maximizing usability. We began by running an evaluation on Amazon Mechanical Turk with over 27,000 respondents to test the usability of different feature combinations. Then we studied user preferences using Google's consumer survey infrastructure. Finally, drawing on the insights gleaned during those studies, we tested our new captcha schemes first on Mechanical Turk and then on a fraction of production traffic. The resulting scheme is now an integral part of our production system and is served to millions of users. Our scheme achieved a 95.3% human accuracy, a 6.7.
Elie Bursztein, Angelique Moscicki, Celine Fabry, Steven Bethard, John C. Mitchell, Daniel Jurafsky
CHI6
2014 Learning to Reason Pragmatically with Cognitive Limitations
Adam Vogel, Andrés Goméz Emilsson, Michael C. Frank, Daniel Jurafsky, Christopher Potts
CogSci4
2014 How to Ask for a Favor: A Case Study on the Success of Altruistic Requests
Tim Althoff, Cristian Danescu-Niculescu-Mizil, Daniel Jurafsky
ICWSM3
2014 On the Importance of Text Analysis for Stock Price Prediction
Heeyoung Lee 0004, Mihai Surdeanu, Bill MacCartney, Daniel Jurafsky
LREC4
2014 Event Extraction Using Distant Supervision
Kevin Reschke, Martin Jankowiak, Mihai Surdeanu, Christopher D. Manning, Daniel Jurafsky
LREC5
2014 Speaker-independent detection of child-directed speech
abstract
Identifying the distinct register that adults use when speaking to children is an important task for child development research. We present a fully automatic, speaker-independent system that detects child-directed speech. The two-stage system uses diarization-style voice activation techniques to extract speech segments followed by a supervised ν-SVM classifier trained on 1582 prosodic and log Mel energy features. The system significantly improves the state of the art, detecting child-directed speech with F1 of .66 (exact boundary) and .83 (within 1 second). A feature analysis confirms the importance of F0 features (especially 3rd quartile and range) as well as new features like the variance, kurtosis, and min of log Mel energy within a frequency band.
Sebastian Schuster 0001, Stephanie Pancoast, Milind Ganjoo, Michael C. Frank, Daniel Jurafsky
SLT5
2014 Charles J. Fillmore
abstract
Charles J. Fillmore died at his home in San Francisco on February 13, 2014, of brain cancer. He was 84 years old. Fillmore was one of the world's pre-eminent scholars of lexical meaning and its relationship with context, grammar, corpora, and computation, and his work had an enormous impact on computational linguistics. His early theoretical work in the 1960s, 1970s, and 1980s on case grammar and then frame semantics significantly influenced computational linguistics, AI, and knowledge representation. More recent work in the last two decades on FrameNet, a computational lexicon and annotated corpus, influenced corpus linguistics and computational lexicography, and led to modern natural language understanding tasks like semantic role labeling.Fillmore was born and raised in St. Paul, Minnesota, and studied linguistics at the University of Minnesota. As an undergraduate he worked on a pre-computational Latin corpus linguistics project, alphabetizing index cards and building concordances. During his service in the Army in the early 1950s he was stationed for three years in Japan. After his service he became the first US soldier to be discharged locally in Japan, and stayed for three years studying Japanese. He supported himself by teaching English, pioneering a way to make ends meet that afterwards became popular with generations of young Americans abroad. In 1957 he moved back to the United States to attend graduate school at the University of Michigan.At Michigan, Fillmore worked on phonetics, phonology, and syntax, first in the American Structuralist tradition of developing what were called “discovery procedures” for linguistic analysis, algorithms for inducing phones or parts of speech. Discovery procedures were thought of as a methodological tool, a formal procedure that linguists could apply to data to discover linguistic structure, for example inducing parts of speech from the slots in “sentence frames” informed by the distribution of surrounding words. Like many linguistic graduate students of the period, he also worked partly on machine translation, and was interviewed at the time by Yehoshua Bar-Hillel, who was touring US machine translation laboratories in preparation for his famous report on the state of MT (Bar-Hillel 1960).Early in his graduate career, however, Fillmore read Noam Chomsky's Syntactic Structures and became an immediate proponent of the new transformational grammar. He graduated with his PhD in 1962 and moved to the linguistics department at Ohio State University. In his early work there Fillmore developed a number of early formal properties of generative grammar, such as the idea that rules would re-apply to representations in iterative stages called cycles (Fillmore 1963), a formal mechanism that still plays a role in modern theories of generative grammar.But his greatest impact on computational linguistics came from the line of research that began with his early work on case grammar (Fillmore 1966, 1968, 1971, 1977a). Fillmore had become interested in argument structure by studying Lucien Tesnière's groundbreaking Éléments de Syntaxe Structurale (Tesnière 1959) in which the term ‘dependency’ was introduced and the foundations were laid for dependency grammar. Like many transformational grammarians of the time, Fillmore began by trying to capture the relationships between distinct formal patterns with systematically related meanings; and he became interested in the different ways of expressing the object and recipient of transfer in sentences like “He gave a book to me” and “He gave me a book” (Fillmore 1962, 1965), a phenomenon that became known as dative movement. He then expanded to the more general goal of representing how the participants in an event are expressed syntactically, as in these two sentences about an event of opening: a. The janitor will open the door with this keyb. This key will open the doorFillmore noticed that despite the differing syntactic structure, in both sentences key plays the role of the instrument of the action and door the role of the object, patient, or theme, and suggested that such abstract roles could constitute a shallow level of meaning representation. Following Tesnière's terminology, Fillmore first referred to these argument roles as actants (Fillmore 1966) but quickly switched to the term case, (see Fillmore (2003)) and proposed a universal list of semantic roles or cases (Agent, Patient, Instrument, etc.), that could be taken on by the arguments of predicates. Verbs would be listed in the lexicon with their ‘case frame’, the list of obligatory (or optional) case arguments.The idea that semantic roles could provide an intermediate level of semantic representation that could help map from syntactic parse structures to deeper, more fully-specified representations of meaning was quickly adopted in natural language processing, and systems for extracting case frames were created for machine translation (Wilks 1973), question-answering (Hendrix, Thompson, and Slocum 1973), spoken-language understanding (Nash-Webber 1975), and dialogue systems (Bobrow et al. 1977). General-purpose semantic role labelers were developed to map to case representations via ATNs (Simmons 1973) or, from parse trees, by using dictionaries with verb-specific case frames (Levin 1977; Marcus 1980). By 1977 case representation was widely used and taught in natural language processing and artificial intelligence, and was described as a standard component of natural language understanding in the first edition of Winston's (1977) textbook Artificial Intelligence.In 1971 Fillmore joined the linguistics faculty at the University of California, Berkeley, and by the mid-1970s he began to expand his ideas on case. He arrived at a more general model of semantic representation, one that expressed the background contexts or perspectives by which a word or a case role could be defined. He called this new representation a frame, and later described the intuition as follows:“The idea behind frame semantics is that speakers are aware of possibly quite complex situation types, packages of connected expectations, that go by various names—frames, schemas, scenarios, scripts, cultural narratives, memes—and the words in our language are understood with such frames as their presupposed background.” (Fillmore 2012, p. 712)He described the name as coming from “the pre-transformationalist view of sentence structure as consisting of a frame and a substitution list,” but the word frame seemed to be in the air for a suite of related notions proposed at about the same time by Minsky (1974), Hymes (1974), and Goffman (1974), as well as related notions with other names like scripts (Schank and Abelson 1975) and schemata (Bobrow and Norman 1975) (see Tannen [1979] for a comparison). Fillmore was also influenced by the semantic field theorists and by a visit to the Yale AI lab where he took notice of the lists of slots and fillers used by early information extraction systems like DeJong (1982) and Schank and Abelson (1977).Fillmore's version of this new idea—more linguistic than other manifestations, focusing on the way that words are associated with frames—was expressed in a series of papers starting in the mid-1970's (Fillmore 1975a, 1976, 1977b, 1982, 1985). His motivating example was the Commercial Event frame, in which a seller sells goods to a buyer, the buyer thus buying the goods that cost a certain amount by paying a price charged by the seller. The definition of each of these verbs (buy, sell, cost, pay, charge), is interrelated by virtue of their joint association with a single kind of event or scenario. The meaning of each word draws in the entire frame, and by using (or hearing) the word, a language user necessarily activates the entire frame. As Fillmore put it:If I tell you that I bought a new pair of shoes, you do not know where I bought them or how much they cost, but you know, by virtue of the frame I have introduced into our discourse, that there have got to be answers to those questions. (Fillmore 1976, p. 29)Fillmore also emphasized the way that frames could represent perspectives on events, such that verbs like sell or pay emphasize different aspects of the same event, or that the differences between alternative senses of the same word might come from their drawing on different frames. Fillmore's linguistic interpretation of frames influenced work in artificial intelligence on knowledge representation like KRL (Bobrow and Winograd 1977), and the perspective-taking aspect of frames had a strong influence on work on framing in linguistics and politics (Lakoff 2010).In 1988 Fillmore taught at the computational linguistics summer school in Pisa run by the late Antonio Zampolli and met the lexicographer Beryl T. Atkins. The two began a collaboration to produce a frame description for the verb risk based on corpus evidence (Fillmore and Atkins 1992). This work, including an invited talk at ACL 1991 (Fillmore and Atkins 1991), influenced the development of other projects in corpus-based lexical semantics (Kipper, Dang, and Palmer 2000; Kipper et al. 2008).Fillmore became interested in this idea that corpus linguistics, lexicography, and lexical semantics could fruitfully be combined (Fillmore 1992) and when he officially retired from UC Berkeley in 1995 he moved to the International Computer Science Institute (ICSI) in Berkeley (although still teaching at UC Berkeley part-time) and began work on the FrameNet project of computational corpus lexicography that combined his early ideas on semantic roles with his later work on frames and his recent interest in corpus lexicography.The idea of FrameNet was to build a large set of frames, each of which consisted of lists of constitutive roles or “frame elements”: sets of words that evoke the frame, grammatical information expressing how each frame element is realized in the sentence, and semantic relations between frames and between frame elements. Corpora were annotated with the evoking words, frames, and frame elements (Baker, Fillmore, and Lowe 1998; Fillmore, Johnson, and Petruck 2003; Fillmore and Baker 2009).Over the next 20 years until his death, Fillmore and his students and colleagues, especially under the direction of Collin Baker, proceeded to create the frames and hand-annotate the corpora. This period of his career was a productive and enjoyable one for Fillmore. In an interview for the ICSI Newsletter, he said“The happiest time of my career has been here at ICSI, where FrameNet has made it possible for me to work with a team of bright young people on a continuing basis doing work that I'll never lose interest in.”The combination of rich linguistic annotation and corpus-based approach instantiated in FrameNet, together with the PropBank semantic-role-labeled corpus created soon afterwards by Martha Palmer and colleagues (Palmer, Kingsbury, and Gildea 2005), led to a revival of automatic approaches to semantic role labeling, first on FrameNet (Gildea and Jurafsky 2000) and then on PropBank data (Gildea and Palmer 2002, inter alia). The problem first addressed in the 1970s by hand-written rules was thus now generally recast as one of supervised machine learning. The resulting plethora of systems for performing automatic semantic role labeling (see the surveys in Palmer, Gildea, and Xue (2010) and Màrquez et al. (2008)) have been applied widely to improve the state of the art in tasks across NLP such as question answering (Shen and Lapata 2007; Surdeanu, Ciaramita, and Zaragoza 2011) and machine translation (Liu and Gildea 2010; Lo et al. 2013). Fillmore's FrameNet project also led to the development of FrameNets for many other languages including Spanish, German, Japanese, Portuguese, Italian, and Chinese. And in a perhaps appropriate return to the discovery procedures that first inspired Fillmore, modern work has focused on ways to induce semantic roles from corpora without role annotation (Swier and Stevenson 2004; Chambers and Jurafsky 2009, 2011; Lang and Lapata 2014).In addition to his work in semantics, Fillmore had significant contributions to syntax and pragmatics, including the influential Santa Cruz Lectures on Deixis (Fillmore 1975b) and a long-standing research project in developing Construction Grammar, a theory—or perhaps more accurately family of theories—that represented a grammar as a collection of constructions, pairings of meaning, and form (Fillmore, Kay, and O'Connor 1988). He also contributed to the application of linguistics to other disciplines including cognitive science, education, and law. Ackerman, Kay, and O'Connor (2014) offer more discussion of these aspects of Fillmore's work.Fillmore was much honored during his career; he was a fellow of the American Academy of Arts and Sciences, served as president of the Linguistic Society of America, was awarded an honorary doctorate from the University of Chicago, had festschrifts and conferences in his honor, received the ACL lifetime achievement award in 2012 (see the text of his acceptance speech in Fillmore [2012]) and, together with Collin Baker, the Antonio Zampolli Prize from ELRA in 2012. Nonetheless, he was unpretentious (universally referred to even by his undergraduates as “Chuck”), modest, embarrassed by compliments, and generally referred to himself light-heartedly as an Ordinary Working Linguist. His Minnesota background (he was Norwegian on his mother's side) always led to Lake Wobegon comparisons, especially given his often bemused smile and wry deadpan wit. His colleague George Lakoff tells the story: “When he first came to Berkeley in 1971, he encountered a culture defined by the then-commonplace expression, ‘Let it all hang out.’ His response was to wear a button saying, ‘Tuck it all back in.’”Fillmore was also a favorite teacher and mentor who enjoyed working with what he often capitalized as “Young People”; and was deeply respected for his brilliance, careful attention to detail, and encyclopedic knowledge of language, and universally beloved for his warmth, generosity, and patience. He is survived by his beloved wife Lily Wong Fillmore, a retired Berkeley linguist and Education professor, their children and grandchildren, and a wide community of fond former colleagues, students, and collaborators, among whom I am proud to include myself.
Daniel Jurafsky
Comput. Linguistics1
2013 A computational approach to politeness with application to social factors
Cristian Danescu-Niculescu-Mizil, Moritz Sudhof, Daniel Jurafsky, Jure Leskovec, Christopher Potts
ACL (1)3
2013 Linguistic Models for Analyzing and Detecting Biased Language
Marta Recasens, Cristian Danescu-Niculescu-Mizil, Daniel Jurafsky
ACL (1)3
2013 Breaking Out of Local Optima with Count Transforms and Model Recombination: A Study in Grammar Induction
abstract
Many statistical learning problems in NLP call for local model search methods.But accuracy tends to suffer with current techniques, which often explore either too narrowly or too broadly: hill-climbers can get stuck in local optima, whereas samplers may be inefficient.We propose to arrange individual local optimizers into organized networks.Our building blocks are operators of two types: (i) transform, which suggests new places to search, via non-random restarts from already-found local optima; and (ii) join, which merges candidate solutions to find better optima.Experiments on grammar induction show that pursuing different transforms (e.g., discarding parts of a learned model or ignoring portions of training data) results in improvements.Groups of locally-optimal solutions can be further perturbed jointly, by constructing mixtures.Using these tools, we designed several modular dependency grammar induction networks of increasing complexity.Our complete system achieves 48.6% accuracy (directed dependency macro-average over all 19 languages in the 2006/7 CoNLL data) -more than 5% higher than the previous state-of-the-art.
Valentin I. Spitkovsky, Hiyan Alshawi, Daniel Jurafsky
EMNLP3
2013 Same Referent, Different Words: Unsupervised Mining of Opaque Coreferent Mentions
Marta Recasens, Matthew Can, Daniel Jurafsky
HLT-NAACL3
2013 Emergence of Gricean Maxims from Multi-Agent Decision Theory
Adam Vogel, Max Bodoia, Christopher Potts, Daniel Jurafsky
HLT-NAACL4
2013 No country for old members: user lifecycle and linguistic change in online communities
abstract
Vibrant online communities are in constant flux. As members join and depart, the interactional norms evolve, stimulating further changes to the membership and its social dynamics. Linguistic change --- in the sense of innovation that becomes accepted as the norm --- is essential to this dynamic process: it both facilitates individual expression and fosters the emergence of a collective identity.
Cristian Danescu-Niculescu-Mizil, Robert West 0001, Daniel Jurafsky, Jure Leskovec, Christopher Potts
WWW3
2013 Deterministic Coreference Resolution Based on Entity-Centric, Precision-Ranked Rules
abstract
We propose a new deterministic approach to coreference resolution that combines the global information and precise features of modern machine-learning models with the transparency and modularity of deterministic, rule-based systems. Our sieve architecture applies a battery of deterministic coreference models one at a time from highest to lowest precision, where each model builds on the previous model's cluster output. The two stages of our sieve-based architecture, a mention detection stage that heavily favors recall, followed by coreference sieves that are precision-oriented, offer a powerful way to achieve both high precision and high recall. Further, our approach makes use of global information through an entity-centric model that encourages the sharing of features across all mentions that point to the same real-world entity. Despite its simplicity, our approach gives state-of-the-art performance on several corpora and genres, and has also been incorporated into hybrid state-of-the-art coreference systems for Chinese and Arabic. Our system thus offers a new paradigm for combining knowledge in rule-based systems that has implications throughout computational linguistics.
Heeyoung Lee 0004, Angel X. Chang, Yves Peirsman, Nathanael Chambers, Mihai Surdeanu, Daniel Jurafsky
Comput. Linguistics6
2013 Detecting friendly, flirtatious, awkward, and assertive speech in speed-dates
Rajesh Ranganath, Daniel Jurafsky, Daniel A. McFarland
Comput. Speech Lang.2
2012 Joint Entity and Event Coreference Resolution across Documents
Heeyoung Lee 0004, Marta Recasens, Angel X. Chang, Mihai Surdeanu, Daniel Jurafsky
EMNLP-CoNLL5
2012 Three Dependency-and-Boundary Models for Grammar Induction
Valentin I. Spitkovsky, Hiyan Alshawi, Daniel Jurafsky
EMNLP-CoNLL3
2012 Learning Attitudes and Attributes from Multi-aspect Reviews
abstract
Most online reviews consist of plain-text feedback together with a single numeric score. However, understanding the multiple `aspects' that contribute to users' ratings may help us to better understand their individual preferences. For example, a user's impression of an audio book presumably depends on aspects such as the story and the narrator, and knowing their opinions on these aspects may help us to recommend better products. In this paper, we build models for rating systems in which such dimensions are explicit, in the sense that users leave separate ratings for each aspect of a product. By introducing new corpora consisting of five million reviews, rated with between three and six aspects, we evaluate our models on three prediction tasks: First, we uncover which parts of a review discuss which of the rated aspects. Second, we summarize reviews by finding the sentences that best explain a user's rating. Finally, since aspect ratings are optional in many of the datasets we consider, we recover ratings that are missing from a user's evaluation. Our model matches state-of-the-art approaches on existing small-scale datasets, while scaling to the real-world datasets we introduce. Moreover, our model is able to `disentangle' content and sentiment words: we automatically learn content words that are indicative of a particular aspect as well as the aspect-specific sentiment words that are indicative of a particular rating.
Julian J. McAuley, Jure Leskovec, Daniel Jurafsky
ICDM3
2012 Learning the Central Events and Participants in Unlabeled Text
Nathanael Chambers, Daniel Jurafsky
ICML2
2012 Parsing Time: Learning to Interpret Time Expressions
Gabor Angeli, Christopher D. Manning, Daniel Jurafsky
HLT-NAACL3
2012 Citation-based bootstrapping for large-scale author disambiguation
abstract
We present a new, two‐stage, self‐supervised algorithm for author disambiguation in large bibliographic databases. In the first “bootstrap” stage, a collection of high‐precision features is used to bootstrap a training set with positive and negative examples of coreferring authors. A supervised feature‐based classifier is then trained on the bootstrap clusters and used to cluster the authors in a larger unlabeled dataset. Our self‐supervised approach shares the advantages of unsupervised approaches (no need for expensive hand labels) as well as supervised approaches (a rich set of features that can be discriminatively trained). The algorithm disambiguates 54,000,000 author instances in Thomson Reuters' Web of Knowledge with B3 F1 of.807. We analyze parameters and features, particularly those from citation networks, which have not been deeply investigated in author disambiguation. The most important citation feature is self‐citation, which can be approximated without expensive extraction of the full network. For the supervised stage, the minor improvement due to other citation features (increasing F1 from.748 to.767) suggests they may not be worth the trouble of extracting from databases that don't already have them. A lean feature set without expensive abstract and title features performs 130 times faster with about equal F1.
Michael Levin 0004, Stefan Krawczyk, Steven Bethard, Daniel Jurafsky
J. Assoc. Inf. Sci. Technol.4
2011 Template-Based Information Extraction without the Templates
Nathanael Chambers, Daniel Jurafsky
ACL2
2011 Punctuation: Making a Point in Unsupervised Dependency Parsing
Valentin I. Spitkovsky, Hiyan Alshawi, Daniel Jurafsky
CoNLL3
2011 Unsupervised Dependency Parsing without Gold Part-of-Speech Tags
Valentin I. Spitkovsky, Hiyan Alshawi, Angel X. Chang, Daniel Jurafsky
EMNLP4
2011 Lateen EM: Unsupervised Training with Multiple Objectives, Applied to Dependency Grammar Induction
Valentin I. Spitkovsky, Hiyan Alshawi, Daniel Jurafsky
EMNLP3
2011 LeadLag LDA: Estimating Topic Specific Leads and Lags of Information Outlets
Ramesh Nallapati, Daniel A. McFarland, Jure Leskovec, Daniel Jurafsky
ICWSM5
2011 Sex, food, and words: the hidden meanings behind everyday language
abstract
Language is a subtle and powerful tool for communication. But the words we use also provide a rich mine of information for the social scientist. The history of words like "ketchup", "ceviche", or "dessert" tells us about the relationships between the superpowers who dominated the globe 500 or 1000 years ago. The words on the back of potato chip packages can demonstrate popular attitudes toward social class. And the names we give ice cream flavors may be an evolutionary reflex of the attempt by early mammals to appear larger than their competitors. The language of dating is just as informative as the language of food. In experiments with speed dating, work in our lab shows that we can detect flirtation or other stances in men and women on dates, just by looking at linguistic features like their pitch, their use of negative words like "can't" or "don't", or how often they use hedges like "sort of" or "kind of". The language of these two popular topics of conversation, food and dating, can teach us a lot about history, culture, and psychology.
Daniel Jurafsky
UIST1
2010 Improving the Use of Pseudo-Words for Evaluating Selectional Preferences
Nathanael Chambers, Daniel Jurafsky
ACL2
2010 Profiting from Mark-Up: Hyper-Text Annotations for Guided Parsing
Valentin I. Spitkovsky, Daniel Jurafsky, Hiyan Alshawi
ACL2
2010 Learning to Follow Navigational Directions
Adam Vogel, Daniel Jurafsky
ACL2
2010 Who should I cite: learning literature search models from citation behavior
abstract
Scientists depend on literature search to find prior work that is relevant to their research ideas. We introduce a retrieval model for literature search that incorporates a wide variety of factors important to researchers, and learns the weights of each of these factors by observing citation patterns. We introduce features like topical similarity and author behavioral patterns, and combine these with features from related work like citation count and recency of publication. We present an iterative process for learning weights for these features that alternates between retrieving articles with the current retrieval model, and updating model weights by training a supervised classifier on these articles. We propose a new task for evaluating the resulting retrieval models, where the retrieval system takes only an abstract as its input and must produce as output the list of references at the end of the abstract's article. We evaluate our model on a collection of journal, conference and workshop articles from the ACL Anthology Reference Corpus. Our model achieves a mean average precision of 28.7, a 12.8 point improvement over a term similarity baseline, and a significant improvement both over models using only features from related work and over models without our iterative learning.
Steven Bethard, Daniel Jurafsky
CIKM2
2010 Viterbi Training Improves Unsupervised Dependency Parsing
Valentin I. Spitkovsky, Hiyan Alshawi, Daniel Jurafsky, Christopher D. Manning
CoNLL3
2010 A Multi-Pass Sieve for Coreference Resolution
Karthik Raghunathan, Heeyoung Lee 0004, Sudarshan Rangarajan, Nathanael Chambers, Mihai Surdeanu, Daniel Jurafsky, Christopher D. Manning
EMNLP6
2010 Parsing to Stanford Dependencies: Trade-offs between Speed and Accuracy
Daniel M. Cer, Marie-Catherine de Marneffe, Daniel Jurafsky, Christopher D. Manning
LREC3
2010 A Database of Narrative Schemas
Nathanael Chambers, Daniel Jurafsky
LREC2
2010 The Best Lexical Metric for Phrase-Based Statistical MT System Optimization
Daniel M. Cer, Christopher D. Manning, Daniel Jurafsky
HLT-NAACL3
2010 From Baby Steps to Leapfrog: How "Less is More" in Unsupervised Dependency Parsing
Valentin I. Spitkovsky, Hiyan Alshawi, Daniel Jurafsky
HLT-NAACL3
2010 How Good Are Humans at Solving CAPTCHAs? A Large Scale Evaluation
abstract
Captchas are designed to be easy for humans but hard for machines. However, most recent research has focused only on making them hard for machines. In this paper, we present what is to the best of our knowledge the first large scale evaluation of captchas from the human perspective, with the goal of assessing how much friction captchas present to the average user. For the purpose of this study we have asked workers from Amazon's Mechanical Turk and an underground captchabreaking service to solve more than 318 000 captchas issued from the 21 most popular captcha schemes (13 images schemes and 8 audio scheme). Analysis of the resulting data reveals that captchas are often difficult for humans, with audio captchas being particularly problematic. We also find some demographic trends indicating, for example, that non-native speakers of English are slower in general and less accurate on English-centric captcha schemes. Evidence from a week's worth of eBay captchas (14,000,000 samples) suggests that the solving accuracies found in our study are close to real-world values, and that improving audio captchas should become a priority, as nearly 1% of all captchas are delivered as audio rather than images. Finally our study also reveals that it is more effective for an attacker to use Mechanical Turk to solve captchas than an underground service.
Elie Bursztein, Steven Bethard, Celine Fabry, John C. Mitchell, Daniel Jurafsky
IEEE Symposium on Security and Privacy5
2010 Which words are hard to recognize? Prosodic, lexical, and disfluency factors that increase speech recognition error rates
Sharon Goldwater, Daniel Jurafsky, Christopher D. Manning
Speech Commun.2
2009 Unsupervised Learning of Narrative Schemas and their Participants
Nathanael Chambers, Daniel Jurafsky
ACL/IJCNLP2
2009 Distant supervision for relation extraction without labeled data
Mike Mintz, Steven Bills, Rion Snow, Daniel Jurafsky
ACL/IJCNLP4
2009 Robust Machine Translation Evaluation with Entailment Features
Sebastian Padó, Michel Galley, Daniel Jurafsky, Christopher D. Manning
ACL/IJCNLP3
2009 It's not you, it's me: Automatically extracting social meaning from speed dates
abstract
Summary form only given. Automatically detecting human social intentions from spoken conversation is an important task for social computing and for dialogue systems. We describe a system for detecting elements of interactional style: whether a speaker is awkward, friendly, or flirtatious. We create and use a new spoken corpus of 991 4-minute speed-dates. Participants rated themselves and each other for these elements of style. Using rich dialogue, lexical, and prosodic features, we are able to detect flirtatious, awkward, and friendly styles in noisy natural conversational data with above 70% accuracy, significantly outperforming not only the baseline but also, for flirtation, outperforming the human interlocutors. We find that features like pitch, energy, and the use of emotional vocabulary help detect flirtation, collaborative conversational style (laughter, questions, collaborative completions) help in detecting friendliness, and disfluencies help in detecting awkwardness. In analyzing why our system outperforms humans, we show that humans are very poor perceivers of flirtatiousness in this task, and instead often project their own intended behavior onto their interlocutors. This talk describes joint work with Dan McFarland (School of Education) and Rajesh Ranganath (Computer Science Department).
Daniel Jurafsky
ASRU1
2009 Hidden Conditional Random Fields for phone recognition
abstract
We apply Hidden Conditional Random Fields (HCRFs) to the task of TIMIT phone recognition. HCRFs are discriminatively trained sequence models that augment conditional random fields with hidden states that are capable of representing subphones and mixture components. We extend HCRFs, which had previously only been applied to phone classification with known boundaries, to recognize continuous phone sequences. We use an N-best inference algorithm in both learning (to approximate all competitor phone sequences) and decoding (to marginalize over hidden states). Our monophone HCRFs achieve 28.3% phone error rate, outperforming maximum likelihood trained HMMs by 3.6%, maximum mutual information trained HMMs by 2.5%, and minimum phone error trained HMMs by 2.2%. We show that this win is partially due to HCRFs' ability to simultaneously optimize discriminative language models and acoustic models, a powerful property that has important implications for speech recognition.
Yun-Hsuan Sung, Daniel Jurafsky
ASRU2
2009 It's Not You, it's Me: Detecting Flirting and its Misperception in Speed-Dates
Rajesh Ranganath, Daniel Jurafsky, Daniel A. McFarland
EMNLP2
2009 Extracting Social Meaning: Identifying Interactional Style in Spoken Conversation
Daniel Jurafsky, Rajesh Ranganath, Daniel A. McFarland
HLT-NAACL1
2009 Measuring machine translation quality as semantic equivalence: A metric based on entailment features
Sebastian Padó, Daniel M. Cer, Michel Galley, Daniel Jurafsky, Christopher D. Manning
Mach. Transl.4
2008 Unsupervised Learning of Narrative Event Chains
Nathanael Chambers, Daniel Jurafsky
ACL2
2008 Which Words Are Hard to Recognize? Prosodic, Lexical, and Disfluency Factors that Increase ASR Error Rates
Sharon Goldwater, Daniel Jurafsky, Christopher D. Manning
ACL2
2008 Jointly Combining Implicit Constraints Improves Temporal Ordering
Nathanael Chambers, Daniel Jurafsky
EMNLP2
2008 Studying the History of Ideas Using Topic Models
David Hall 0006, Daniel Jurafsky, Christopher D. Manning
EMNLP2
2008 Cheap and Fast - But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks
Rion Snow, Brendan T. O'Connor 0001, Daniel Jurafsky, Andrew Y. Ng
EMNLP3
2008 Maximum conditional likelihood linear regression and maximum a posteriori for hidden conditional random fields speaker adaptation
abstract
This paper shows how to improve Hidden Conditional Random Fields (HCRFs) for phone classification by applying various speaker adaptation techniques. These include Maximum A Posteriori (MAP) adaptation as well as a new technique we introduce called Maximum Conditional Likelihood Linear Regression (MCLLR), a discriminative variant of the widely used MLLR algorithm. In previous work, we and others have shown that HCRFs outperform even discriminatively trained HMMs. In this paper we show that HCRFs adapted via MCLLR or via MAP adaptation also work better than similarly adapted HMMs. We also compare MCLLR and MAP adaptation performance with different amounts of adaptation data. MCLLR adaptation performs better when the amount of adaptation data is relatively small, while MAP adaptation outperforms MCLLR with larger amounts of adaptation.
Yun-Hsuan Sung, Constantinos Boulis, Daniel Jurafsky
ICASSP3
2007 Classifying Temporal Relations Between Events
Nathanael Chambers, Shan Wang 0002, Daniel Jurafsky
ACL3
2007 Measuring Importance and Query Relevance in Topic-focused Multi-document Summarization
Ani Nenkova, Daniel Jurafsky
ACL3
2007 Disambiguating Between Generic and Referential "You" in Dialog
Matthew Purver, Daniel Jurafsky
ACL3
2007 Automatic detection of contrastive elements in spontaneous speech
abstract
In natural speech people use different levels of prominence to signal which parts of an utterance are especially important. Contrastive elements are often produced with stronger than usual prominence and their presence modifies the meaning of the utterance in subtle but important ways. We use a richly annotated corpus of conversational speech to study the acoustic characteristics of contrastive elements and the differences between them and words at other levels of prominence. We report our results for automatic detection of contrastive elements based on acoustic and textual features, finding that a baseline predicting nouns and adjectives as contrastive performs on par with the best combination of features. We achieve a much better performance in a modified task of detecting contrastive elements among words that are predicted to bear pitch accent.
Ani Nenkova, Daniel Jurafsky
ASRU2
2007 Regularization, adaptation, and non-independent features improve hidden conditional random fields for phone classification
abstract
We show a number of improvements in the use of Hidden Conditional Random Fields (HCRFs) for phone classification on the TIMIT and Switchboard corpora. We first show that the use of regularization effectively prevents overfitting, improving over other methods such as early stopping. We then show that HCRFs are able to make use of non-independent features in phone classification, at least with small numbers of mixture components, while HMMs degrade due to their strong independence assumptions. Finally, we successfully apply Maximum a Posteriori adaptation to HCRFs, decreasing the phone classification error rate in the Switchboard corpus by around 1% – 5% given only small amounts of adaptation data.
Yun-Hsuan Sung, Constantinos Boulis, Christopher D. Manning, Daniel Jurafsky
ASRU4
2007 Learning to Merge Word Senses
Rion Snow, Sushant Prakash, Daniel Jurafsky, Andrew Y. Ng
EMNLP-CoNLL3
2007 Modelling prominence and emphasis improves unit-selection synthesis
abstract
We describe the results of large scale perception experiments showing improvements in synthesising two distinct kinds of prominence: standard pitch-accent and strong emphatic accents. Previously prominence assignment has been mainly evaluated by computing accuracy on a prominence-labelled test set. By contrast we integrated an automatic pitch-accent classifier into the unit selection target cost and showed that listeners preferred these synthesised sentences. We also describe an improved recording script for collecting emphatic accents, and show that generating emphatic accents leads to further improvements in the fiction genre over incorporating pitch accent only. Finally, we show differences in the effects of prominence between child-directed speech and news and fiction genres. Index Terms: speech synthesis, prosody, prominence, pitch accent, unit selection
Volker Strom, Ani Nenkova, Robert A. J. Clark, Yolanda Vazquez-Alvarez, Jason M. Brenier, Simon King 0001, Daniel Jurafsky
INTERSPEECH7
2007 To Memorize or to Predict: Prominence labeling in Conversational Speech
Ani Nenkova, Jason M. Brenier, Anubha Kothari, Sasha Calhoun, Laura Whitton, David Beaver, Daniel Jurafsky
HLT-NAACL7
2006 Semantic Taxonomy Induction from Heterogenous Evidence
abstract
We propose a novel algorithm for inducing semantic taxonomies. Previous algorithms for taxonomy induction have typically focused on independent classifiers for discovering new single relationships based on hand-constructed or automatically discovered textual patterns. By contrast, our algorithm flexibly incorporates evidence from multiple classifiers over heterogenous relationships to optimize the entire structure of the taxonomy, using knowledge of a word's coordinate terms to help in determining its hypernyms, and vice versa. We apply our algorithm on the problem of sense-disambiguated noun hyponym acquisition, where we combine the predictions of hypernym and coordinate term classifiers with the knowledge in a preexisting semantic taxonomy (WordNet 2.1). We add 10,000 novel synsets to WordNet 2.1 at 84% precision, a relative error reduction of 70% over a non-joint algorithm using the same component classifiers. Finally, we show that a taxonomy built using our algorithm shows a 23% relative F-score improvement over WordNet 2.1 on an independent testset of hypernym pairs.
Rion Snow, Daniel Jurafsky, Andrew Y. Ng
ACL2
2006 Detection of word fragments in Mandarin telephone conversation
abstract
We describe preliminary work on the detection of word fragments in Mandarin conversational telephone speech. We extracted prosodic, voice quality, and lexical features, and trained Decision Tree and SVM classifiers. Previous research shows that glottalization features are instrumental in English fragment detection. However, we show that Mandarin fragments are quite different than English; 90% of Mandarin fragments are followed immediately by a repetition of the fragmentary word. These repetition fragments are not glottalized, and they have a very specific distribution; the 12 most frequent words (“you”, “I”, “that”, “have”, “then”, etc.) cover 50% of the tokens of these fragments. Thus rather than glottalization, we found the most useful feature for Mandarin fragment detection was the identity of the neighboring character (word or morpheme). In an oracle experiment using the true (reference) neighboring words as well as prosodic and voice quality features, we achieved 80% accuracy in Mandarin fragment detection.
Cheng-Tao Chu, Yun-Hsuan Sung, Daniel Jurafsky
INTERSPEECH4
2006 Limitations of MLLR adaptation with Spanish-accented English: an error analysis
abstract
We studied the effect of MLLR adaptation with Spanishaccented English to understand the strengths and weaknesses of MLLR with unseen foreign accents. We trained a global MLLR transform on 10 adaptation sentences per speaker, giving a 3.4 % absolute decrease in phone error rate. We then studied the pattern of improvements across phones and phone classes. Phones that improved the least tended to be those that do not exist in Spanish. Results suggest the poorer performance is related to increased insertion and substituter rates during the adaptation phase, as well as greater acoustic variability. Index Terms: speech recognition, adaptation, MLLR, foreign accent, Spanish, error analysis
Constance Clarke, Daniel Jurafsky
INTERSPEECH2
2006 Have we met? MDP based speaker ID for robot dialogue
abstract
We present a novel approach to speaker identification in robot dialogue that allows a robot to recognize people during natural conversation and address them by name. Our STanford AI Robot (STAIR) dialogue system attempts to mirror the human speaker identification process. We model the robot dialogue problem as a Markov Decision Process (MDP) and apply a reinforcement learning algorithm to try to learn the optimal dialogue actions. The MDP model works in conjunction with a traditional statistical cluster based speaker identification subsystem. Our approach also addresses open-set speaker identification, dynamically adding new speaker profiles as well as continuously updating known profiles. Index Terms: dialogue, MDP, speaker identification, speaker recognition, robot conversation
Filip Krsmanovic, Curtis Spencer, Daniel Jurafsky, Andrew Y. Ng
INTERSPEECH3
2006 The (Non)Utility of Linguistic Features for Predicting prominence in spontaneous speech
abstract
Conversational speech is characterized by prosodic variability which makes pitch accent prediction for this genre especially difficult. The linguistic literature points out that complex features such as information status, contrast and animacy help predict pitch accent placement. In this paper, we use a corpus annotated for such features to determine if they improve prominence prediction over traditional shallow features such as frequency and part-of-speech, or over new ones that we introduce. We demonstrate that while correlated with prominence, complex linguistic features do not improve prediction accuracy. Furthermore, the performance of our classifier is quite close to the ceiling defined by variability in human accent placement. An oracle experiment demonstrates, though, that at least some accuracy improvement is still possible.
Jason M. Brenier, Ani Nenkova, Anubha Kothari, Laura Whitton, David Beaver, Daniel Jurafsky
SLT6
2006 A Dialectal Chinese Speech Recognition Framework
Thomas Zheng, William J. Byrne, Daniel Jurafsky
J. Comput. Sci. Technol.4
2005 Semantic Role Labeling Using Different Syntactic Views
abstract
Semantic role labeling is the process of annotating the predicate-argument structure in text with semantic labels. In this paper we present a state-of-the-art baseline semantic role labeling system based on Support Vector Machine classifiers. We show improvements on this system by: i) adding new features including features extracted from dependency parses, ii) performing feature selection and calibration and iii) combining parses obtained from semantic parsers trained using different syntactic views. Error analysis of the baseline system showed that approximately half of the argument identification errors resulted from parse errors in which there was no syntactic constituent that aligned with the correct argument. In order to address this problem, we combined semantic parses from a Minipar syntactic parse and from a chunked syntactic representation with our original baseline system which was based on Charniak parses. All of the reported techniques resulted in performance improvements.
Sameer Pradhan, Wayne H. Ward, Kadri Hacioglu, James H. Martin, Daniel Jurafsky
ACL5
2005 Semantic Role Chunking Combining Complementary Syntactic Views
Sameer Pradhan, Kadri Hacioglu, Wayne H. Ward, James H. Martin, Daniel Jurafsky
CoNLL5
2005 The detection of emphatic words using acoustic and lexical features
abstract
In this study, we describe an automatic detector for prosodically salient or emphasized words in speech. Knowledge of whether a word is emphatic or not could improve Text-to-Speech synthesis as well as spoken language summarization. Previous work on emphasis detection has focused on the automatic recognition of pitch accents. Our model extends earlier research by automatically identifying emphatic pitch accents, a subset of pitch accents that mark special discourse functions with extreme degrees of salience. The overall best performance achieved by our system was 87.8 % correct, 8.0 % above baseline performance. The results of a feature selection algorithm show that the top-performing features in our models are primarily acoustic measures. Our work identifies important cues for emphasis in speech and shows that it is possible for an automated system to distinguish between two levels of perceived prominence in pitch accents with a high degree of accuracy. 1.
Jason M. Brenier, Daniel M. Cer, Daniel Jurafsky
INTERSPEECH3
2005 Pitch accent prediction: effects of genre and speaker
abstract
To build a robust pitch accent prediction system, we need to understand the effects of speech genre and speaker variation. This paper reports our studies on genre and speaker variation in pitch accent placement and their effects on automatic pitch accent prediction. We find some interesting accentuation pattern differences that can be attributed to speech genre, and a set of textual features that are robust to genre in accent prediction. We also find that although there is significant variation among speakers in pitch accent placement, speaker dependent models are not needed in accent prediction. Finally, we show that after taking speaker variation into account, there is little room to improve for state-of-the-art classifiers on read news speech. 1.
Jiahong Yuan, Jason M. Brenier, Daniel Jurafsky
INTERSPEECH3
2005 Accent detection and speech recognition for Shanghai-accented Mandarin
abstract
As speech recognition systems are used in ever more applications, it is crucial for the systems to be able to deal with accented speakers. Various techniques, such as acoustic model adaptation and pronunciation adaptation, have been reported to improve the recognition of non-native or accented speech. In this paper, we propose a new approach that combines accent detection, accent discriminative acoustic features, acoustic adaptation and model selection for accented Chinese speech recognition. Experimental results show that this approach can improve the recognition of accented speech. 1.
Yanli Zheng, Richard Sproat, Liang Gu, Izhak Shafran, Haolang Zhou, Daniel Jurafsky, Rebecca Starr, Su-Youn Yoon
INTERSPEECH7
2005 Support Vector Learning for Semantic Argument Classification
Sameer Pradhan, Kadri Hacioglu, Valerie Krugler, Wayne H. Ward, James H. Martin, Daniel Jurafsky
Mach. Learn.6
2005 Editorial
Eric Fosler-Lussier, William J. Byrne, Daniel Jurafsky
Speech Commun.3
2004 Semantic Role Labeling by Tagging Syntactic Chunks
Kadri Hacioglu, Sameer Pradhan, Wayne H. Ward, James H. Martin, Daniel Jurafsky
CoNLL5
2004 Shallow Semantic Parsing using Support Vector Machines
Sameer Pradhan, Wayne H. Ward, Kadri Hacioglu, James H. Martin, Daniel Jurafsky
HLT-NAACL5
2004 Shallow Semantc Parsing of Chinese
Honglin Sun, Daniel Jurafsky
HLT-NAACL2
2004 Learning Syntactic Patterns for Automatic Hypernym Discovery
abstract
Semantic taxonomies such as WordNet provide a rich source of knowl- edge for natural language processing applications, but are expensive to build, maintain, and extend. Motivated by the problem of automatically constructing and extending such taxonomies, in this paper we present a new algorithm for automatically learning hypernym (is-a) relations from text. Our method generalizes earlier work that had relied on using small numbers of hand-crafted regular expression patterns to identify hyper- nym pairs. Using "dependency path" features extracted from parse trees, we introduce a general-purpose formalization and generalization of these patterns. Given a training set of text containing known hypernym pairs, our algorithm automatically extracts useful dependency paths and applies them to new corpora to identify novel pairs. On our evaluation task (de- termining whether two nouns in a news article participate in a hypernym relationship), our automatically extracted database of hypernyms attains both higher precision and higher recall than WordNet.
Rion Snow, Daniel Jurafsky, Andrew Y. Ng
NIPS2
2003 Semantic Role Parsing: Adding Semantic Structure to Unstructured Text
abstract
There is an ever-growing need to add structure in the form of semantic markup to the huge amounts of unstructured text data now available. We present the technique of shallow semantic parsing, the process of assigning a simple WHO did WHAT to WHOM, etc., structure to sentences in text, as a useful tool in achieving this goal. We formulate the semantic parsing problem as a classification problem using support vector machines. Using a hand-labeled training set and a set of features drawn from earlier work together with some feature enhancements, we demonstrate a system that performs better than all other published results on shallow semantic parsing.
Sameer Pradhan, Kadri Hacioglu, Wayne H. Ward, James H. Martin, Daniel Jurafsky
ICDM5
2002 Automatic Labeling of Semantic Roles
abstract
We present a system for identifying the semantic relationships, or semantic roles, filled by constituents of a sentence within a semantic frame. Given an input sentence and a target word and frame, the system labels constituents with either abstract semantic roles, such as Agent or Patient, or more domain-specific semantic roles, such as Speaker, Message, and Topic. The system is based on statistical classifiers trained on roughly 50,000 sentences that were hand-annotated with semantic roles by the FrameNet semantic labeling project. We then parsed each training sentence into a syntactic tree and extracted various lexical and syntactic features, including the phrase type of each constituent, its grammatical function, and its position in the sentence. These features were combined with knowledge of the predicate verb, noun, or adjective, as well as information such as the prior probabilities of various combinations of semantic roles. We used various lexical clustering algorithms to generalize across possible fillers of roles. Test sentences were parsed, were annotated with these features, and were then passed through the classifiers. Our system achieves 82% accuracy in identifying the semantic role of presegmented constituents. At the more difficult task of simultaneously segmenting constituents and identifying their semantic role, the system achieved 65% precision and 61% recall. Our study also allowed us to compare the usefulness of different features and feature combination methods in the semantic role labeling task. We also explore the integration of role labeling with statistical syntactic parsing and attempt to generalize to predicates unseen in the training data.
Daniel Gildea, Daniel Jurafsky
Comput. Linguistics2
2001 Is Knowledge-Free Induction of Multiword Unit Dictionary Headwords a Solved Problem?
Patrick Schone, Daniel Jurafsky
EMNLP2
2001 The effect of language model probability on pronunciation reduction
abstract
We investigate how the probability of a word affects its pronunciation. We examined 5618 tokens of the 10 most frequent (function) words in Switchboard and 2042 tokens of content words whose lexical form ends in a t or d. Our observations were drawn from the phonetically hand-transcribed subset of the Switchboard corpus, enabling us to code each word with its pronunciation and duration. Using linear and logistic regression to control for contextual factors, we show that words which have a high unigram, bigram, or reverse bigram (given the following word) probability are shorter, more likely to have a reduced vowel, and more likely to have a deleted final t or d. These results suggest that pronunciation models in speech recognition and synthesis should take into account word probability given both the previous and following words, for both content and function words.
Daniel Jurafsky, Alan Bell, Michelle Gregory, William D. Raymond
ICASSP1
2001 What kind of pronunciation variation is hard for triphones to model?
abstract
In order to help understand why gains in pronunciation modeling have proven so elusive, we investigated which kinds of pronunciation variation are well captured by triphone models, and which are not. We do this by examining the change in behavior of a recognizer as it receives further triphone training. We show that many of the kinds of variation which previous pronunciation models attempt to capture, including phone substitution or phone reduction, are in fact already well captured by triphones. Our analysis suggests new areas where future pronunciation models should focus, including syllable deletion.
Daniel Jurafsky, Wayne H. Ward, Zhang Banping, Keith Herold, Xiuyang Yu, Zhang Sen
ICASSP1
2001 Knowledge-Free Induction of Inflectional Morphologies
Patrick Schone, Daniel Jurafsky
NAACL2
2001 A Bayesian Model Predicts Human Parse Preference and Reading Times in Sentence Processing
abstract
Narayanan and Jurafsky (1998) proposed that human language compre- hension can be modeled by treating human comprehenders as Bayesian reasoners, and modeling the comprehension process with Bayesian de- cision trees. In this paper we extend the Narayanan and Jurafsky model to make further predictions about reading time given the probability of difference parses or interpretations, and test the model against reading time data from a psycholinguistic experiment.
Daniel Jurafsky
NIPS2
2000 Automatic Labeling of Semantic Roles
abstract
We present a system for identifying the semantic relationships, or semantic roles, filled by constituents of a sentence within a semantic frame. Various lexical and syntactic features are derived from parse trees and used to derive statistical classifiers from hand-annotated training data.
Daniel Gildea, Daniel Jurafsky
ACL2
2000 Dialog Act Modeling for Automatic Tagging and Recognition of Conversational Speech
abstract
We describe a statistical approach for modeling dialogue acts in conversational speech, i.e., speech-act-like units such as STATEMENT, Question, BACKCHANNEL, Agreement, Disagreement, and Apology. Our model detects and predicts dialogue acts based on lexical, collocational, and prosodic cues, as well as on the discourse coherence of the dialogue act sequence. The dialogue model is based on treating the discourse structure of a conversation as a hidden Markov model and the individual dialogue acts as observations emanating from the model states. Constraints on the likely sequence of dialogue acts are modeled via a dialogue act n-gram. The statistical dialogue grammar is combined with word n-grams, decision trees, and neural networks modeling the idiosyncratic lexical and prosodic manifestations of each dialogue act. We develop a probabilistic integration of speech recognition with dialogue modeling, to improve both speech recognition and dialogue act classification accuracy. Models are trained and evaluated using a large hand-labeled database of 1,155 conversations from the Switchboard corpus of spontaneous human-to-human telephone speech. We achieved good dialogue act labeling accuracy (65% based on errorful, automatically recognized words and prosody, and 71% based on word transcripts, compared to a chance baseline accuracy of 35% and human accuracy of 84%) and a small reduction in word recognition error.
Andreas Stolcke, Klaus Ries 0001, Noah Coccaro, Elizabeth Shriberg, Rebecca Bates 0001, Daniel Jurafsky, Paul Taylor 0001, Rachel Martin, Carol Van Ess-Dykema, Marie Meteer
Comput. Linguistics6
1998 Towards better integration of semantic predictors in statistical language modeling
abstract
We introduce a number of techniques designed to help integrate semantic knowledge with N-gram language models for automatic speech recognition. Our techniques allow us to integrate Latent Semantic Analysis (LSA), a word-similarity algorithm based on word co-occurrence information, with N-gram models. While LSA is good at predicting content words which are coherent with the rest of a text, it is a bad predictor of frequent words, has a low dynamic range, and is inaccurate when combined linearly with N-grams. We show that modifying the dynamic range, applying a per-word confidence metric, and using geometric rather than linear combinations with N-grams produces a more robust language model which has a lower perplexity on a Wall Street Journal testset than a baseline N-gram model. 1. INTRODUCTION There has been a lot of recent work on augmenting n-gram language models with other information sources such as longer distance syntactic, and semantic constraints (e.g. [8], [6]). In previous ...
Noah Coccaro, Daniel Jurafsky
ICSLP2
1998 Reduction of English function words in switchboard
abstract
The causes of pronunciation reduction in 8458 occurrences of ten frequent English function words in a four-hour sample from conversations from the Switchboard corpus were examined. Using ordinary linear and logistic regression models, we examined the length of the words, the form of their vowel (basic, full, or reduced) , and final obstruent deletion. For all of these we found strong, independent effects of speaking rate, predictability, the form of the following word, and planning problem disfluencies. The results bear on issues in speech recognition, models of speech production, and conversational analysis. 1. INTRODUCTION This study reports the results of an investigation of some factors affecting the reduction or lenition of ten of the most frequent English words, namely I, and, the, that, a, you, to, of, it, and in, in the Switchboard corpus of conversational speech. Frequent function words are of particular interest because they are not only subject to the contextual and stylis...
Daniel Jurafsky, Alan Bell, Eric Fosler-Lussier, Cynthia Girand, William D. Raymond
ICSLP1
1998 An American national corpus: a proposal
Charles J. Fillmore, Nancy Ide, Daniel Jurafsky, Catherine Macleod
LREC3
1996 Learning Bias and Phonological-Rule Induction
Daniel Gildea, Daniel Jurafsky
Comput. Linguistics2
1995 Automatic Induction of Finite State Transducers for Simple Phonological Rules
abstract
This paper presents a method for learning phonological rules from sample pairs of underlying and surface forms, without negative evidence. The learned rules are represented as finite state transducers that accept underlying forms as input and generate surface forms as output. The algorithm for learning them is an extension of the OSTIA algorithm for learning general subsequential finite state transducers. Although OSTIA is capable of learning arbitrary s.f.s.t's in the limit, large dictionaries of actual English pronunciations did not give enough samples to correctly induce phonological rules. We then augmented OSTIA with two kinds of knowledge specific to natural language phonology, biases from "universal grammar". One bias is that underlying phones are often realized as phonetically similar or identical surface phones. The other biases phonological rules to apply across natural phonological classes. The additions helped in learning more compact, accurate, and general transducers than the unmodified OSTIA algorithm. An implementation of the algorithm successfully learns a number of English postlexical rules.
Daniel Gildea, Daniel Jurafsky
ACL2
1995 Learning Phonological Rule Probabilities from Speech Corpora with Exploratory Computational Phonology
abstract
This paper presents an algorithm for learning the probabilities of optional phonological rules from corpora. The algorithm is based on using a speech recognition system to discover the surface pronunciations of words in speech corpora; using an automatic system obviates expensive phonetic labeling by hand. We describe the details of our algorithm and show the probabilities the system has learned for ten common phonological rules which model reductions and coarticulation effects. These probabilities were derived from a corpus of 7203 sentences of read speech from the Wall Street Journal, and are shown to be a reasonably close match to probabilities from phonetically hand-transcribed data (TIMIT). Finally, we analyze the probability differences between rule use in male versus female speech, and suggest that the differences are caused by differing average rates of speech.
Gary N. Tajchman, Daniel Jurafsky, Eric Fosler-Lussier
ACL2
1995 Using a stochastic context-free grammar as a language model for speech recognition
abstract
This paper describes a number of experiments in adding new grammatical knowledge to the Berkeley Restaurant Project (BeRP), our medium-vocabulary (1300 word), speaker-independent, spontaneous continuous-speech understanding system. We describe an algorithm for using a probabilistic Earley parser and a stochastic context-free grammar (SCFG) to generate word transition probabilities at each frame for a Viterbi decoder. We show that using an SCFG as a language model improves the word error rate from 34.6% (bigram) to 29.6% (SCFG), and the semantic sentence recognition error from from 39.0% (bigram) to 34.1% (SCFG). In addition, we get a further reduction to 28.8% word error by mixing the bigram and SCFG LMs. We also report on our preliminary results from using discourse-context information in the LM.
Daniel Jurafsky, Chuck Wooters, Jonathan Segal, Andreas Stolcke, Eric Fosler-Lussier, Gary N. Tajchman, Nelson Morgan
ICASSP1
1995 Building multiple pronunciation models for novel words using exploratory computational phonology
abstract
In this paper we describe a completely automatic algorithm that builds multiple pronunciation word models by expanding baseform pronunciations with a set of candidate phonological rules. We show how to train the probabilities of these phonological rules, and how to use these probabilities to assign pronunciation probabilities to words not seen in the training corpus. The algorithm we propose is an instance of the class of techniques we call Exploratory Computational Phonology. 1. INTRODUCTION One well-known difficulty in understanding speakerindependent continuous speech is variability in the pronunciation of words. This variability occurs across speakers and also across different contexts for a single speaker. In order to model this variation, recognition systems often use a richer lexicon in which each word has multiple pronunciations. Using a multiple-pronunciation lexicon requires setting a probability for each pronunciation. The minimal algorithm, for example, would assign each ...
Gary N. Tajchman, Eric Fosler-Lussier, Daniel Jurafsky
EUROSPEECH3
1994 The berkeley restaurant project
abstract
This paper describes the architecture and performance of the Berkeley Restaurant Project (BeRP), a medium-vocabulary, speaker-independent, spontaneous continuous speech understanding system currently under development at ICSI. BeRP serves as a testbed for a number of our speech-related research projects, including robust feature extraction, connectionist phonetic likelihood estimation, automatic induction of multiplepronunciation lexicons, foreign accent detection and modeling, advanced language models, and lip-reading. In addition, it has proved quite usable in its function as a database frontend, even though many of our subjects are non-native speakers of English.
Daniel Jurafsky, Chuck Wooters, Gary N. Tajchman, Jonathan Segal, Andreas Stolcke, Eric Fosler-Lussier, Nelson Morgan
ICSLP1
1992 An On-Line Computational Model of Human Sentence Interpretation
Daniel Jurafsky
AAAI1
1990 Representing and Integrating Linguistic Knowledge
Daniel Jurafsky
COLING1
1989 James Allen, Understanding Natural Language
Daniel Jurafsky
Artif. Intell.1
1988 Issues in Relating Syntax Semantics
Daniel Jurafsky
COLING1