Kathy McKeown

dblp:m/KathleenMcKeown · also Kathleen McKeown, Kathleen R. McKeown · DBLP profile ↗
← Back
195ranked-venue papers
27as first author
53since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 173 · 19 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 9 first-author · 2 since 2021Databases, data management, data science and information retrieval · 9 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 9 · 2 first-authorHuman-computer interaction and ubiquitous computing · 4 · 1 first-authorComputer networks · 1Security and privacy · 1
YearPublicationVenuePosition
2026 Characterizing and Evaluating Working Emotion Vocabularies in Multilingual Large Language Models
abstract
Prior work evaluating emotion and affective understanding in large language models (LLMs) typically rely on predetermined label sets or focus on a singular evaluation task (e.g., emotion detection).We consider affective states, referring to the much broader variety of terms people use to label their emotional experiences.We evaluate multilingual language models' understanding of affective states in English and Spanish through three different tasks: 1) identification, where models predict an affective state given text, 2) expression, where models generate text expressing a given affective state, and 3) verification, where models report whether a given term refers to an affective state.We show that performance on one task is not necessarily predictive of performance on another.Using these three tasks, we then begin to explore when and why models struggle to understand particular affective states compared to others.We examine systematic patterns in the affective state terms that are well and poorly understood by models, characterizing the working emotion vocabulary of LLMs.
Nicholas Deas, Iván Pérez Mejía, Ellie Yang, Kathy McKeown
ACL (1)4
2025 Enhancing Multimodal Affective Analysis with Learned Live Comment Features
abstract
Live comments, also known as Danmaku, are user-generated messages that are synchronized with video content. These comments overlay directly onto streaming videos, capturing viewer emotions and reactions in real-time. While prior work has leveraged live comments in affective analysis, its use has been limited due to the relative rarity of live comments across different video platforms. To address this, we first construct the Live Comment for Affective Analysis (LCAffect) dataset, which contains live comments for English and Chinese videos spanning diverse genres that elicit a wide spectrum of emotions. Then, using this dataset, we use contrastive learning to train a video encoder to produce synthetic live comment features for enhanced multimodal affective content analysis. Through comprehensive experimentation on a wide range of affective analysis tasks (sentiment, emotion recognition, and sarcasm detection) in both English and Chinese, we demonstrate that these synthetic live comment features significantly improve performance over state-of-the-art methods.
Zhaoyuan Deng, Amith Ananthram, Kathy McKeown
AAAI3
2025 Data Caricatures: On the Representation of African American Language in Pretraining Corpora
abstract
Nicholas Deas, Blake Vente, Amith Ananthram, Jessica A Grieser, Desmond U. Patton, Shana Kleiner, James R. Shepard Iii, Kathleen McKeown. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Nicholas Deas, Blake Vente, Amith Ananthram, Jessica Grieser, Desmond Upton Patton, Shana Kleiner, James Shepard, Kathy McKeown
ACL (1)8
2025 Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution
abstract
Recent state-of-the-art authorship attribution methods learn authorship representations of text in a latent, uninterpretable space, which hinders their usability in real-world applications. We propose a novel approach for interpreting learned embeddings by identifying representative points in the latent space and leveraging large language models to generate informative natural language descriptions of the writing style associated with each point. We evaluate the alignment between our interpretable and latent spaces and demonstrate superior prediction agreement over baseline methods. Additionally, we conduct a human evaluation to assess the quality of these style descriptions and validate their utility in explaining the latent space. Finally, we show that human performance on the challenging authorship attribution task improves by +20% on average when aided with explanations from our method.
Milad Alshomary, Narutatsu Ri, Marianna Apidianaki, Ajay Patel, Smaranda Muresan, Kathy McKeown
COLING6
2025 Summarization of Opinionated Political Documents with Varied Perspectives
abstract
Global partisan hostility and polarization has increased, and this polarization is heightened around presidential elections. Models capable of generating accurate summaries of diverse perspectives can help reduce such polarization by exposing users to alternative perspectives. In this work, we introduce a novel dataset and task for independently summarizing each political perspective in a set of passages from opinionated news articles. For this task, we propose a framework for evaluating different dimensions of perspective summary performance. We benchmark 11 summarization models and LLMs of varying sizes and architectures through both automatic and human evaluation. While recent models like GPT-4o perform well on this task, we find that all models struggle to generate summaries that are faithful to the intended perspective. Our analysis of summaries focuses on how extraction behavior is impacted by features of the input documents.
Nicholas Deas, Kathy McKeown
COLING2
2025 ManiTweet: A New Benchmark for Identifying Manipulation of News on Social Media
abstract
Considerable advancements have been made to tackle the misrepresentation of information derived from reference articles in the domains of fact-checking and faithful summarization. However, an unaddressed aspect remains - the identification of social media posts that manipulate information within associated news articles. This task presents a significant challenge, primarily due to the prevalence of personal opinions in such posts. We present a novel task, identifying manipulation of news on social media, which aims to detect manipulation in social media posts and identify manipulated or inserted information. To study this task, we have proposed a data collection schema and curated a dataset called ManiTweet, consisting of 3.6K pairs of tweets and corresponding articles. Our analysis demonstrates that this task is highly challenging, with large language models (LLMs) yielding unsatisfactory performance. Additionally, we have developed a simple yet effective basic model that outperforms LLMs significantly on the ManiTweet dataset. Finally, we have conducted an exploratory analysis of human-written tweets, unveiling intriguing connections between manipulation and the domain and factuality of news articles, as well as revealing that manipulated sentences are more likely to encapsulate the main story or consequences of a news outlet.
Kung-Hsiang Huang, Hou Pong Chan, Kathy McKeown, Heng Ji 0001
COLING3
2025 Layered Insights: Generalizable Analysis of Human Authorial Style by Leveraging All Transformer Layers
abstract
We propose a new approach for the authorship attribution task that leverages the various linguistic representations learned at different layers of pre-trained transformer-based models.We evaluate our approach on two popular authorship attribution models and three evaluation datasets, in in-domain and out-of-domain scenarios.We found that utilizing various transformer layers improves the robustness of authorship attribution models when tested on outof-domain data, resulting in a much stronger performance.Our analysis gives further insights into how our model's different layers get specialized in representing certain linguistic aspects that we believe benefit the model when tested out of the domain.
Milad Alshomary, Nikhil Reddy Varimalla, Vishal Anand 0002, Smaranda Muresan, Kathy McKeown
EMNLP5
2025 Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions
abstract
We introduce and study artificial impressionspatterns in LLMs' internal representations of prompts that resemble human impressions and stereotypes based on language.We fit linear probes on generated prompts to predict impressions according to the two-dimensional Stereotype Content Model (SCM).Using these probes, we study the relationship between impressions and downstream model behavior as well as prompt features that may inform such impressions.We find that LLMs inconsistently report impressions when prompted, but also that impressions are more consistently linearly decodable from their hidden representations.Additionally, we show that artificial impressions of prompts are predictive of the quality and use of hedging in model responses.We also investigate how particular content, stylistic, and dialectal features in prompts impact LLM impressions.
Nicholas Deas, Kathy McKeown
EMNLP2
2025 Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding
abstract
Determining faithfulness of a claim to a source document is an important problem across many domains.This task is generally treated as a binary judgment of whether the claim is supported or unsupported in relation to the source.In many cases, though, whether a claim is supported can be ambiguous.For instance, it may depend on making inferences from given evidence, and different people can reasonably interpret the claim as either supported or unsupported based on their agreement with those inferences.Forcing binary labels upon such claims lowers the reliability of evaluation.In this work, we reframe the task to manage the subjectivity involved with factuality judgments of ambiguous claims.We introduce LLMgenerated edits of summaries as a method of providing a nuanced evaluation of claims: how much does a summary need to be edited to be unambiguous?Whether a claim gets rewritten and how much it changes can be used as an automatic evaluation metric, the Ambiguity Rewrite Metric (ARM), with a much richer feedback signal than a binary judgment of faithfulness.We focus on the area of narrative summarization as it is particularly rife with ambiguity and subjective interpretation.We show that ARM produces a 21% absolute improvement in annotator agreement on claim faithfulness, indicating that subjectivity is reduced.
Melanie Subbiah, Akankshya Mishra, Grace Kim, Liyan Tang, Greg Durrett, Kathy McKeown
EMNLP6
2025 Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
abstract
Large Language Models (LLMs) are typically trained to reflect a relatively uniform set of values, which limits their applicability to tasks that require understanding of nuanced human perspectives.Recent research has underscored the importance of enabling LLMs to support steerable pluralism -the capacity to adopt a specific perspective and align generated outputs with it.In this work, we investigate whether Chain-of-Thought (CoT) reasoning techniques can be applied to building steerable pluralistic models.We explore several methods, including CoT prompting, fine-tuning on humanauthored CoT, fine-tuning on synthetic explanations, and Reinforcement Learning with Verifiable Rewards (RLVR).We evaluate these approaches using the Value Kaleidoscope and OpinionQA datasets.Among the methods studied, RLVR consistently outperforms others and demonstrates strong training sample efficiency.We further analyze the generated CoT traces with respect to faithfulness and safety.
Kathy McKeown, Smaranda Muresan
EMNLP2
2025 See It from My Perspective: How Language Affects Cultural Bias in Image Understanding
abstract
Vision-language models (VLMs) can respond to queries about images in many languages. However, beyond language, culture affects how we see things. For example, individuals from Western cultures focus more on the central figure in an image while individuals from East Asian cultures attend more to scene context (Nisbett 2001). In this work, we characterize the Western bias of VLMs in image understanding and investigate the role that language plays in this disparity. We evaluate VLMs across subjective and objective visual tasks with culturally diverse images and annotations. We find that VLMs perform better on the Western split than on the East Asian split of each task. Through controlled experimentation, we trace one source of this bias in image understanding to the lack of diversity in language model construction. While inference in a language nearer to a culture can lead to reductions in bias, we show it is much more effective when that language was well-represented during text-only pre-training. Interestingly, this yields bias reductions even when prompting in English. Our work highlights the importance of richer representation of all languages in building equitable VLMs.
Amith Ananthram, Elias Stengel-Eskin, Mohit Bansal, Kathy McKeown
ICLR4
2025 A General Framework for Inference-time Scaling and Steering of Diffusion Models
abstract
Diffusion models have demonstrated remarkable performance in generative modeling, but generating samples with specific desiderata remains challenging. Existing solutions --- such as fine-tuning, best-of-n sampling, and gradient-based guidance --- are expensive, inefficient, or limited in applicability. In this work, we propose FK steering, a framework for inference-time steering diffusion models with reward functions. In this work, we introduce FK steering, which applies Feynman-Kac interacting particle systems to the inference-time steering of diffusion models with arbitrary reward functions. FK steering works by generating multiple trajectories, called particles, and resampling particles at intermediate steps based on scores computed using functions called potentials. Potentials are defined using rewards for intermediate states and are chosen such that a high score indicates the particle will yield a high-reward sample. We explore various choices of potentials, rewards, and samplers. Steering text-to-image models with a human preference reward, we find that FK steering outperforms fine-tuned models with just 2 particles. Moreover, FK steering a 0.8B parameter model outperforms a 2.6B model, achieving state-of-the-art performance on prompt fidelity. We also steer text diffusion models with rewards for text quality and rare attributes such as toxicity, and find that FK steering generates lower perplexity text and enables gradient-free control. Overall, inference-time scaling and steering of diffusion models, even training-free, provides significant quality and controllability benefits. Code available [here](https://github.com/zacharyhorvitz/FK-Diffusion-Steering).
Raghav Singhal, Zachary Horvitz, Ryan Teehan, Mengye Ren, Kathy McKeown, Rajesh Ranganath
ICML6
2025 Counterfactual Simulatability of LLM Explanations for Generation Tasks
abstract
LLMs can be unpredictable, as even slight alterations to the prompt can cause the output to change in unexpected ways. Thus, the ability of models to accurately explain their behavior is critical, especially in high-stakes settings. Counterfactual simulatability measures how well an explanation allows users to infer the model’s output on related counterfactuals and has been previously studied for yes/no question answering. We provide a general framework for extending this method to generation tasks, using news summarization and medical suggestion as example use cases. We find that while LLM explanations do enable users to better predict their outputs on counterfactuals in the summarization setting, there is significant room for improvement for medical suggestion. Furthermore, our results suggest that evaluating counterfactual simulatability may be more appropriate for skill-based tasks as opposed to knowledge-based tasks.
Marvin Limpijankit, Yanda Chen, Melanie Subbiah, Nicholas Deas, Kathy McKeown
INLG5
2025 Mining Contextualized Visual Associations from Images for Creativity Understanding
abstract
Understanding another person’s creative output requires a shared language of association. However, when training vision-language models such as CLIP, we rely on web-scraped datasets containing short, predominantly literal, alt-text. In this work, we introduce a method for mining contextualized associations for salient visual elements in an image that can scale to any unlabeled dataset. Given an image, we can use these mined associations to generate high quality creative captions at increasing degrees of abstraction. With our method, we produce a new dataset of visual associations and 1.7m creative captions for the images in MSCOCO. Human evaluation confirms that these captions remain visually grounded while exhibiting recognizably increasing abstraction. Moreover, fine-tuning a visual encoder on this dataset yields meaningful improvements in zero-shot image-text retrieval in two creative domains: poetry and metaphor visualization. We release our dataset, our generation code and our models for use by the broader community.
Ananya Sahu, Amith Ananthram, Kathy McKeown
INLG3
2025 Forecasting Conversation Derailments Through Generation
abstract
Forecasting conversation derailment can be useful in real-world settings such as online content moderation, conflict resolution, and business negotiations. However, despite language models’ success at identifying offensive speech present in conversations, they struggle to forecast future conversation derailments. In contrast to prior work that predicts conversation outcomes solely based on the past conversation history, our approach samples multiple future conversation trajectories conditioned on existing conversation history using a fine-tuned LLM. It predicts the conversation outcome based on the consensus of these trajectories. We also experimented with leveraging socio-linguistic attributes, which reflect turn-level conversation dynamics, as guidance when generating future conversations. Our method of future conversation trajectories surpasses state-of-the-art results on English conversation derailment prediction benchmarks and demonstrates significant accuracy gains in ablation studies.
Kathy McKeown, Smaranda Muresan
INLG2
2025 Evaluating Defeasible Reasoning in LLMs with DEFREASING
abstract
Emily Allaway, Kathleen McKeown. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Emily Allaway, Kathy McKeown
NAACL (Long Papers)2
2025 StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples
abstract
Ajay Patel, Jiacheng Zhu, Justin Qiu, Zachary Horvitz, Marianna Apidianaki, Kathleen McKeown, Chris Callison-Burch. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ajay Patel, Justin Qiu, Zachary Horvitz, Marianna Apidianaki, Kathy McKeown, Chris Callison-Burch
NAACL (Long Papers)6
2025 Guiding LLM Decision-Making with Fairness Reward Models
abstract
Large language models are increasingly used to support high-stakes decisions, potentially influencing who is granted bail or receives a loan. Naive chain-of-thought sampling can improve average decision accuracy, but has also been shown to amplify unfair bias. To address this challenge and enable the trustworthy use of reasoning models in high-stakes decision-making, we propose a framework for training a generalizable Fairness Reward Model (FRM). Our model assigns a fairness score to LLM reasoning, enabling the system to down-weight biased trajectories and favor equitable ones when aggregating decisions across reasoning chains. We show that a single Fairness Reward Model, trained on weakly supervised, LLM-annotated examples of biased versus unbiased reasoning, transfers across tasks, domains, and model families without additional fine-tuning. When applied to real-world decision-making tasks including recidivism prediction and social media moderation, our approach consistently improves fairness while matching, or even surpassing, baseline accuracy.
Zara Hall, Melanie Subbiah, Thomas P. Zollo, Kathy McKeown, Richard S. Zemel
NeurIPS4
2024 ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer
abstract
Textual style transfer is the task of transforming stylistic properties of text while preserving meaning. Target "styles" can be defined in numerous ways, ranging from single attributes (e.g. formality) to authorship (e.g. Shakespeare). Previous unsupervised style-transfer approaches generally rely on significant amounts of labeled data for only a fixed set of styles or require large language models. In contrast, we introduce a novel diffusion-based framework for general-purpose style transfer that can be flexibly adapted to arbitrary target styles at inference time. Our parameter-efficient approach, ParaGuide, leverages paraphrase-conditioned diffusion models alongside gradient-based guidance from both off-the-shelf classifiers and strong existing style embedders to transform the style of text while preserving semantic information. We validate the method on the Enron Email Corpus, with both human and automatic evaluations, and find that it outperforms strong baselines on formality, sentiment, and even authorship style transfer.
Zachary Horvitz, Ajay Patel, Chris Callison-Burch, Kathy McKeown
AAAI5
2024 Parallel Structures in Pre-training Data Yield In-Context Learning
abstract
Pre-trained language models (LMs) are capable of in-context learning (ICL): they can adapt to a task with only a few examples given in the prompt without any parameter update.However, it is unclear where this capability comes from as there is a stark distribution shift between pre-training text and ICL prompts.In this work, we study what patterns of the pretraining data contribute to ICL.We find that LMs' ICL ability depends on parallel structures in the pre-training data-pairs of phrases following similar templates in the same context window.Specifically, we detect parallel structures by checking whether training on one phrase improves prediction of the other, and conduct ablation experiments to study their effect on ICL.We show that removing parallel structures in the pre-training data reduces LMs' ICL accuracy by 51% (vs 2% from random ablation).This drop persists even when excluding common patterns such as n-gram repetitions and long-range dependency, showing the diversity and generality of parallel structures.A closer look at the detected parallel structures indicates that they cover diverse linguistic tasks and span long distances in the data.
Yanda Chen, Chen Zhao 0013, Kathy McKeown, He He 0001
ACL (1)4
2024 Social Orientation: A New Feature for Dialogue Analysis
abstract
There are many settings where it is useful to predict and explain the success or failure of a dialogue. Circumplex theory from psychology models the social orientations (e.g., Warm-Agreeable, Arrogant-Calculating) of conversation participants and can be used to predict and explain the outcome of social interactions. Our work is novel in its systematic application of social orientation tags to modeling conversation outcomes. In this paper, we introduce a new data set of dialogue utterances machine-labeled with social orientation tags. We show that social orientation tags improve task performance, especially in low-resource settings, on both English and Chinese language benchmarks. We also demonstrate how social orientation tags help explain the outcomes of social interactions when used in neural models. Based on these results showing the utility of social orientation tags for dialogue outcome prediction tasks, we release our data sets, code, and models that are fine-tuned to predict social orientation tags on dialogue utterances.
Todd Morrill, Zhaoyuan Deng, Yanda Chen, Amith Ananthram, Colin Wayne Leach, Kathy McKeown
LREC/COLING6
2024 MASIVE: Open-Ended Affective State Identification in English and Spanish
abstract
In the field of emotion analysis, much NLP research focuses on identifying a limited number of discrete emotion categories, often applied across languages.These basic sets, however, are rarely designed with textual data in mind, and culture, language, and dialect can influence how particular emotions are interpreted.In this work, we broaden our scope to a practically unbounded set of affective states, which includes any terms that humans use to describe their experiences of feeling.We collect and publish MASIVE, a dataset of Reddit posts in English and Spanish containing over 1,000 unique affective states each.We then define the new problem of affective state identification for language generation models framed as a masked span prediction task.On this task, we find that smaller fine-tuned multilingual models outperform much larger LLMs, even on region-specific Spanish affective states.Additionally, we show that pre-training on MA-SIVE improves model performance on existing emotion benchmarks.Finally, through machine translation experiments, we find that native speaker-written data is vital to good performance on this task.
Nicholas Deas, Elsbeth Turcan, Iván Pérez Mejía, Kathy McKeown
EMNLP4
2024 STORYSUMM: Evaluating Faithfulness in Story Summarization
abstract
Human evaluation has been the gold standard for checking faithfulness in abstractive summarization.However, with a challenging source domain like narrative, multiple annotators can agree a summary is faithful, while missing details that are obvious errors only once pointed out.We therefore introduce a new dataset, STORYSUMM, comprising LLM summaries of short stories with localized faithfulness labels and error explanations.This benchmark is for evaluation methods, testing whether a given method can detect challenging inconsistencies.Using this dataset, we first show that any one human annotation protocol is likely to miss inconsistencies, and we advocate for pursuing a range of methods when establishing ground truth for a summarization dataset.We finally test recent automatic metrics and find that none of them achieve more than 70% balanced accuracy on this task, demonstrating that it is a challenging benchmark for future work in faithfulness evaluation.
Melanie Subbiah, Faisal Ladhak, Akankshya Mishra, Griffin Adams, Lydia B. Chilton, Kathy McKeown
EMNLP6
2024 Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations
abstract
Large language models (LLMs) are trained to imitate humans to explain human decisions. However, do LLMs explain themselves? Can they help humans build mental models of how LLMs process different inputs? To answer these questions, we propose to evaluate $\textbf{counterfactual simulatability}$ of natural language explanations: whether an explanation can enable humans to precisely infer the model’s outputs on diverse counterfactuals of the explained input. For example, if a model answers ”$\textit{yes}$” to the input question ”$\textit{Can eagles fly?}$” with the explanation ”$\textit{all birds can fly}$”, then humans would infer from the explanation that it would also answer ”$\textit{yes}$” to the counterfactual input ”$\textit{Can penguins fly?}$”. If the explanation is precise, then the model’s answer should match humans’ expectations. We implemented two metrics based on counterfactual simulatability: precision and generality. We generated diverse counterfactuals automatically using LLMs. We then used these metrics to evaluate state-of-the-art LLMs (e.g., GPT-4) on two tasks: multi-hop factual reasoning and reward modeling. We found that LLM’s explanations have low precision and that precision does not correlate with plausibility. Therefore, naively optimizing human approvals (e.g., RLHF) may be insufficient.
Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 0013, He He 0001, Jacob Steinhardt, Kathy McKeown
ICML8
2024 TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization
abstract
Liyan Tang, Igor Shalyminov, Amy Wong, Jon Burnsky, Jake Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Lijia Sun, Yi Zhang, Saab Mansour, Kathleen McKeown. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Liyan Tang, Igor Shalyminov, Amy Wing-mei Wong, Jon Burnsky, Jake W. Vincent, Siffi Singh, Song Feng 0001, Hwanjun Song, Lijia Sun, Yi Zhang 0053, Saab Mansour, Kathy McKeown
NAACL-HLT14
2024 Fair Abstractive Summarization of Diverse Perspectives
abstract
Yusen Zhang, Nan Zhang, Yixin Liu, Alexander Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao, Dragomir Radev, Kathleen McKeown, Rui Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yusen Zhang 0001, Yixin Liu 0003, Alexander R. Fabbri, Junru Liu, Ryo Kamoi, Xiaoxin Lu, Caiming Xiong, Jieyu Zhao 0001, Dragomir R. Radev, Kathy McKeown, Rui Zhang 0037
NAACL-HLT11
2024 Exceptions, Instantiations, and Overgeneralization: Insights into How Language Models Process Generics
abstract
Abstract Large language models (LLMs) have garnered a great deal of attention for their exceptional generative performance on commonsense and reasoning tasks. In this work, we investigate LLMs’ capabilities for generalization using a particularly challenging type of statement: generics. Generics express generalizations (e.g., birds can fly) but do so without explicit quantification. They are notable because they generalize over their instantiations (e.g., sparrows can fly) yet hold true even in the presence of exceptions (e.g., penguins do not). For humans, these generic generalizations play a fundamental role in cognition, concept acquisition, and intuitive reasoning. We investigate how LLMs respond to and reason about generics. To this end, we first propose a framework grounded in pragmatics to automatically generate both exceptions and instantiations – collectively exemplars. We make use of focus—a pragmatic phenomenon that highlights meaning-bearing elements in a sentence—to capture the full range of interpretations of generics across different contexts of use. This allows us to derive precise logical definitions for exemplars and operationalize them to automatically generate exemplars from LLMs. Using our system, we generate a dataset of ∼370kexemplars across ∼17k generics and conduct a human validation of a sample of the generated data. We use our final generated dataset to investigate how LLMs reason about generics. Humans have a documented tendency to conflate universally quantified statements (e.g., all birds can fly) with generics. Therefore, we probe whether LLMs exhibit similar overgeneralization behavior in terms of quantification and in property inheritance. We find that LLMs do show evidence of overgeneralization, although they sometimes struggle to reason about exceptions. Furthermore, we find that LLMs may exhibit similar non-logical behavior to humans when considering property inheritance from generics.
Emily Allaway, Chandra Bhagavatula, Jena D. Hwang, Kathy McKeown, Sarah-Jane Leslie
Comput. Linguistics4
2024 Reading Subtext: Evaluating Large Language Models on Short Story Summarization with Writers
abstract
Abstract We evaluate recent Large Language Models (LLMs) on the challenging task of summarizing short stories, which can be lengthy, and include nuanced subtext or scrambled timelines. Importantly, we work directly with authors to ensure that the stories have not been shared online (and therefore are unseen by the models), and to obtain informed evaluations of summary quality using judgments from the authors themselves. Through quantitative and qualitative analysis grounded in narrative theory, we compare GPT-4, Claude-2.1, and LLama-2-70B. We find that all three models make faithfulness mistakes in over 50% of summaries and struggle with specificity and interpretation of difficult subtext. We additionally demonstrate that LLM ratings and other automatic metrics for summary quality do not correlate well with the quality ratings from the writers.
Melanie Subbiah, Sean Zhang, Lydia B. Chilton, Kathy McKeown
Trans. Assoc. Comput. Linguistics4
2024 Benchmarking Large Language Models for News Summarization
abstract
Abstract Large language models (LLMs) have shown promise for automatic summarization but the reasons behind their successes are poorly understood. By conducting a human evaluation on ten LLMs across different pretraining methods, prompts, and model scales, we make two important observations. First, we find instruction tuning, not model size, is the key to the LLM’s zero-shot summarization capability. Second, existing studies have been limited by low-quality references, leading to underestimates of human performance and lower few-shot and finetuning performance. To better evaluate LLMs, we perform human evaluation over high-quality summaries we collect from freelance writers. Despite major stylistic differences such as the amount of paraphrasing, we find that LLM summaries are judged to be on par with human written summaries.
Faisal Ladhak, Esin Durmus, Percy Liang, Kathy McKeown, Tatsunori B. Hashimoto
Trans. Assoc. Comput. Linguistics5
2023 Generating EDU Extracts for Plan-Guided Summary Re-Ranking
abstract
Two-step approaches, in which summary candidates are generated-then-reranked to return a single summary, can improve ROUGE scores over the standard single-step approach.Yet, standard decoding methods (i.e., beam search, nucleus sampling, and diverse beam search) produce candidates with redundant, and often low quality, content.In this paper, we design a novel method to generate candidates for re-ranking that addresses these issues.We ground each candidate abstract on its own unique content plan and generate distinct plan-guided abstracts using a model's top beam.More concretely, a standard language model (a BART LM) auto-regressively generates elemental discourse unit (EDU) content plans with an extractive copy mechanism.The top K beams from the content plan generator are then used to guide a separate LM, which produces a single abstractive candidate for each distinct plan.We apply an existing re-ranker (BRIO) to abstractive candidates generated from our method, as well as baseline decoding methods.We show large relevance improvements over previously published methods on widely used single document news article corpora, with ROUGE-2 F1 gains of 0.88, 2.01, and 0.38 on CNN / Dailymail, NYT, and Xsum, respectively.A human evaluation on CNN / DM validates these results.Similarly, on 1k samples from CNN / DM, we show that prompting GPT-3 to follow EDU plans outperforms sampling-based methods by 1.05 ROUGE-2 F1 points.Code to generate and realize plans is available at https: //github.com/griff4692/edu-sum.
Griffin Adams, Alexander R. Fabbri, Faisal Ladhak, Noémie Elhadad, Kathy McKeown
ACL (1)5
2023 Faking Fake News for Real Fake News Detection: Propaganda-Loaded Training Data Generation
abstract
Despite recent advances in detecting fake news generated by neural models, their results are not readily applicable to effective detection of human-written disinformation.What limits the successful transfer between them is the sizable gap between machine-generated fake news and human-authored ones, including the notable differences in terms of style and underlying intent.With this in mind, we propose a novel framework for generating training examples that are informed by the known styles and strategies of human-authored propaganda.Specifically, we perform self-critical sequence training guided by natural language inference to ensure the validity of the generated articles, while also incorporating propaganda techniques, such as appeal to authority and loaded language.In particular, we create a new training dataset, PROPANEWS, with 2,256 examples, which we release for future use.Our experimental results show that fake news detectors trained on PROPANEWS are better at detecting human-written disinformation by 3.62-7.69%F1 score on two public datasets.1
Kung-Hsiang Huang, Kathy McKeown, Preslav Nakov, Yejin Choi 0001, Heng Ji 0001
ACL (1)2
2023 Few-Shot Data-to-Text Generation via Unified Representation and Multi-Source Learning
abstract
Alexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma, Patrick Ng, Zhiguo Wang, Bonan Min, William Yang Wang, Kathleen McKeown, Vittorio Castelli, Dan Roth, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Alexander Hanbo Li, Mingyue Shang, Evangelia Spiliopoulou, Jie Ma 0005, Patrick Ng, Zhiguo Wang 0006, Bonan Min, William Yang Wang, Kathy McKeown, Vittorio Castelli, Dan Roth 0001, Bing Xiang
ACL (1)9
2023 Unsupervised Selective Rationalization with Noise Injection
abstract
A major issue with using deep learning models in sensitive applications is that they provide no explanation for their output.To address this problem, unsupervised selective rationalization produces rationales alongside predictions by chaining two jointly-trained components, a rationale generator and a predictor.Although this architecture guarantees that the prediction relies solely on the rationale, it does not ensure that the rationale contains a plausible explanation for the prediction.We introduce a novel training technique that effectively limits generation of implausible rationales by injecting noise between the generator and the predictor.Furthermore, we propose a new benchmark for evaluating unsupervised selective rationalization models using movie reviews from existing datasets.We achieve sizeable improvements in rationale plausibility and task accuracy over the state-of-the-art across a variety of tasks, including our new benchmark, while maintaining or improving model faithfulness.1
Adam Storek, Melanie Subbiah, Kathy McKeown
ACL (1)3
2023 Penguins Don't Fly: Reasoning about Generics through Instantiations and Exceptions
abstract
Emily Allaway, Jena D. Hwang, Chandra Bhagavatula, Kathleen McKeown, Doug Downey, Yejin Choi. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Emily Allaway, Jena D. Hwang, Chandra Bhagavatula, Kathy McKeown, Doug Downey, Yejin Choi 0001
EACL4
2023 When Do Pre-Training Biases Propagate to Downstream Tasks? A Case Study in Text Summarization
abstract
Faisal Ladhak, Esin Durmus, Mirac Suzgun, Tianyi Zhang, Dan Jurafsky, Kathleen McKeown, Tatsunori Hashimoto. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Faisal Ladhak, Esin Durmus, Mirac Suzgun, Daniel Jurafsky, Kathy McKeown, Tatsunori B. Hashimoto
EACL6
2023 Faithfulness-Aware Decoding Strategies for Abstractive Summarization
abstract
Despite significant progress in understanding and improving faithfulness in abstractive summarization, the question of how decoding strategies affect faithfulness is less studied.We present a systematic study of the effect of generation techniques such as beam search and nucleus sampling on faithfulness in abstractive summarization.We find a consistent trend where beam search with large beam sizes produces the most faithful summaries while nucleus sampling generates the least faithful ones.We propose two faithfulness-aware generation methods to further improve faithfulness over current generation techniques: (1) ranking candidates generated by beam search using automatic faithfulness metrics and (2) incorporating lookahead heuristics that produce a faithfulness score on the future summary.We show that both generation methods significantly improve faithfulness across two datasets as evaluated by four automatic faithfulness metrics and human evaluation.To reduce computational cost, we demonstrate a simple distillation approach that allows the model to generate faithful summaries with just greedy decoding.1
David Wan, Mengwen Liu, Kathy McKeown, Markus Dreyer, Mohit Bansal
EACL3
2023 Evaluation of African American Language Bias in Natural Language Generation
abstract
Warning: This paper contains content and language that may be considered offensive to some readers.While biases disadvantaging African American Language (AAL) have been uncovered in models for tasks such as speech recognition and toxicity detection, there has been little investigation of these biases for language generation models like ChatGPT.We evaluate how well LLMs understand AAL in comparison to White Mainstream English (WME), the encouraged "standard" form of English taught in American classrooms.We measure large language model performance on two tasks: a counterpart generation task, where a model generates AAL given WME and vice versa, as well as a masked span prediction (MSP) task, where models predict a phrase hidden from their input.Using a novel dataset of AAL texts from a variety of regions and contexts, we present evidence of dialectal bias for six pre-trained LLMs through performance gaps on these tasks.
Nicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Upton Patton, Elsbeth Turcan, Kathy McKeown
EMNLP6
2022 Faithful or Extractive? On Mitigating the Faithfulness-Abstractiveness Trade-off in Abstractive Summarization
abstract
Despite recent progress in abstractive summarization, systems still suffer from faithfulness errors.While prior work has proposed models that improve faithfulness, it is unclear whether the improvement comes from an increased level of extractiveness of the model outputs as one naive way to improve faithfulness is to make summarization models more extractive.In this work, we present a framework for evaluating the effective faithfulness of summarization systems, by generating a faithfulnessabstractiveness trade-off curve that serves as a control at different operating points on the abstractiveness spectrum.We then show that the baseline system as well as recently proposed methods for improving faithfulness, fail to consistently improve over the control at the same level of abstractiveness.Finally, we learn a selector to identify the most faithful and abstractive summary for a given document, and show that this system can attain higher faithfulness scores in human evaluations while being more abstractive than the baseline system on two datasets.Moreover, we show that our system is able to achieve a better faithfulnessabstractiveness trade-off than the control at the same level of abstractiveness.
Faisal Ladhak, Esin Durmus, He He 0001, Claire Cardie, Kathy McKeown
ACL (1)5
2022 Using Structured Content Plans for Fine-grained Syntactic Control in Pretrained Language Model Generation
abstract
Large pretrained language models offer powerful generation capabilities, but cannot be reliably controlled at a sub-sentential level. We propose to make such fine-grained control possible in pretrained LMs by generating text directly from a semantic representation, Abstract Meaning Representation (AMR), which is augmented at the node level with syntactic control tags. We experiment with English-language generation of three modes of syntax relevant to the framing of a sentence - verb voice, verb tense, and realization of human entities - and demonstrate that they can be reliably controlled, even in settings that diverge drastically from the training distribution. These syntactic aspects contribute to how information is framed in text, something that is important for applications such as summarization which aim to highlight salient information.
Fei-Tzin Lee, Miguel Ballesteros, Feng Nan, Kathy McKeown
COLING4
2022 Constrained Regeneration for Cross-Lingual Query-Focused Extractive Summarization
abstract
Query-focused summaries of foreign-language, retrieved documents can help a user understand whether a document is actually relevant to the query term. A standard approach to this problem is to first translate the source documents and then perform extractive summarization to find relevant snippets. However, in a cross-lingual setting, the query term does not necessarily appear in the translations of relevant documents. In this work, we show that constrained machine translation and constrained post-editing can improve human relevance judgments by including a query term in a summary when its translation appears in the source document. We also present several strategies for selecting only certain documents for regeneration which yield further improvements
Elsbeth Turcan, David Wan, Faisal Ladhak, Petra Galuscáková, Sukanta Sen, Svetlana Tchistiakova, Weijia Xu, Marine Carpuat, Kenneth Heafield, Douglas W. Oard, Kathy McKeown
COLING11
2022 SafeText: A Benchmark for Exploring Physical Safety in Language Models
abstract
Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia Chilton, Desmond Patton, Kathleen McKeown, William Yang Wang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Sharon Levy, Emily Allaway, Melanie Subbiah, Lydia B. Chilton, Desmond Upton Patton, Kathy McKeown, William Yang Wang
EMNLP6
2021 Cross-language Sentence Selection via Data Augmentation and Rationale Training
abstract
Yanda Chen, Chris Kedzie, Suraj Nair, Petra Galuscakova, Rui Zhang, Douglas Oard, Kathleen McKeown. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yanda Chen, Chris Kedzie, Suraj Nair 0001, Petra Galuscáková, Rui Zhang 0037, Douglas W. Oard, Kathy McKeown
ACL/IJCNLP (1)7
2021 InfoSurgeon: Cross-Media Fine-grained Information Consistency Checking for Fake News Detection
abstract
Yi Fung, Christopher Thomas, Revanth Gangi Reddy, Sandeep Polisetty, Heng Ji, Shih-Fu Chang, Kathleen McKeown, Mohit Bansal, Avi Sil. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yi R. Fung 0001, Christopher Thomas 0004, Revanth Gangi Reddy, Sandeep Polisetty, Heng Ji 0001, Shih-Fu Chang, Kathy McKeown, Mohit Bansal, Avirup Sil
ACL/IJCNLP (1)7
2021 Improving Factual Consistency of Abstractive Summarization via Question Answering
abstract
Feng Nan, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Feng Nan, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathy McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang 0006, Andrew O. Arnold, Bing Xiang
ACL/IJCNLP (1)5
2021 A Unified Feature Representation for Lexical Connotations
abstract
Ideological attitudes and stance are often expressed through subtle meanings of words and phrases.Understanding these connotations is critical to recognizing the cultural and emotional perspectives of the speaker.In this paper, we use distant labeling to create a new lexical resource representing connotation aspects for nouns and adjectives.Our analysis shows that it aligns well with human judgments.Additionally, we present a method for creating lexical representations that capture connotations within the embedding space and show that using the embeddings provides a statistically significant improvement on the task of stance detection when data is limited.
Emily Allaway, Kathy McKeown
EACL2
2021 Entity-level Factual Consistency of Abstractive Text Summarization
abstract
Feng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, Bing Xiang. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Feng Nan, Ramesh Nallapati, Zhiguo Wang 0006, Cícero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathy McKeown, Bing Xiang
EACL7
2021 Event-Driven News Stream Clustering using Entity-Aware Contextual Embeddings
abstract
Kailash Karthik Saravanakumar, Miguel Ballesteros, Muthu Kumar Chandrasekaran, Kathleen McKeown. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Kailash Karthik Saravanakumar, Miguel Ballesteros, Muthu Kumar Chandrasekaran, Kathy McKeown
EACL4
2021 Segmenting Subtitles for Correcting ASR Segmentation Errors
abstract
David Wan, Chris Kedzie, Faisal Ladhak, Elsbeth Turcan, Petra Galuscakova, Elena Zotkina, Zhengping Jiang, Peter Bell, Kathleen McKeown. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
David Wan, Chris Kedzie, Faisal Ladhak, Elsbeth Turcan, Petra Galuscáková, Elena Zotkina, Zhengping Jiang, Peter Bell 0001, Kathy McKeown
EACL9
2021 A Bag of Tricks for Dialogue Summarization
abstract
Dialogue summarization comes with its own peculiar challenges as opposed to news or scientific articles summarization.In this work, we explore four different challenges of the task: handling and differentiating parts of the dialogue belonging to multiple speakers, negation understanding, reasoning about the situation, and informal language understanding.Using a pretrained sequence-to-sequence language model, we explore speaker name substitution, negation scope highlighting, multi-task learning with relevant tasks, and pretraining on in-domain data.Our experiments show that our proposed techniques indeed improve summarization performance, outperforming strong baselines.
Muhammad Khalifa, Miguel Ballesteros, Kathy McKeown
EMNLP (1)3
2021 Timeline Summarization based on Event Graph Compression via Time-Aware Optimal Transport
abstract
Timeline Summarization identifies major events from a news collection and describes them following temporal order, with key dates tagged.Previous methods generally generate summaries separately for each date after they determine the key dates of events.These methods overlook the events' intra-structures (arguments) and inter-structures (event-event connections).Following a different route, we propose to represent the news articles as an event-graph, thus the summarization task becomes compressing the whole graph to its salient sub-graph.The key hypothesis is that the events connected through shared arguments and temporal order depict the skeleton of a timeline, containing events that are semantically related, structurally salient, and temporally coherent in the global event graph.A time-aware optimal transport distance is then introduced for learning the compression model in an unsupervised manner.We show that our approach significantly improves the state of the art on three real-world datasets, including two public standard benchmarks and our newly collected Timeline 100 dataset. 1
Manling Li, Tengfei Ma 0001, Mo Yu, Lingfei Wu 0001, Tian Gao 0007, Heng Ji 0001, Kathy McKeown
EMNLP (1)7
2021 Adversarial Learning for Zero-Shot Stance Detection on Social Media
abstract
Emily Allaway, Malavika Srikanth, Kathleen McKeown. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Emily Allaway, Malavika Srikanth, Kathy McKeown
NAACL-HLT3
2021 Emotion-Infused Models for Explainable Psychological Stress Detection
abstract
Elsbeth Turcan, Smaranda Muresan, Kathleen McKeown. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Elsbeth Turcan, Smaranda Muresan, Kathy McKeown
NAACL-HLT3
2021 Supporting Clustering with Contrastive Learning
abstract
Dejiao Zhang, Feng Nan, Xiaokai Wei, Shang-Wen Li, Henghui Zhu, Kathleen McKeown, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Dejiao Zhang, Feng Nan, Xiaokai Wei, Shang-Wen Li 0001, Henghui Zhu, Kathy McKeown, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang
NAACL-HLT6
2020 Exploring Content Selection in Summarization of Novel Chapters
abstract
We present a new summarization task, generating summaries of novel chapters using summary/chapter pairs from online study guides.This is a harder task than the news summarization task, given the chapter length as well as the extreme paraphrasing and generalization found in the summaries.We focus on extractive summarization, which requires the creation of a gold-standard set of extractive summaries.We present a new metric for aligning reference summary sentences with chapter sentences to create gold extracts and also experiment with different alignment methods.Our experiments demonstrate significant improvement over prior alignment approaches for our task as shown through automatic metrics and a crowd-sourced pyramid analysis.
Faisal Ladhak, Bryan Li, Yaser Al-Onaizan, Kathy McKeown
ACL4
2020 Contextual Analysis of Social Media: The Promise and Challenge of Eliciting Context in Social Media Posts with Natural Language Processing
abstract
While natural language processing affords researchers an opportunity to automatically scan millions of social media posts, there is growing concern that automated computational tools lack the ability to understand context and nuance in human communication and language. This article introduces a critical systematic approach for extracting culture, context and nuance in social media data. The Contextual Analysis of Social Media (CASM) approach considers and critiques the gap between inadequacies in natural language processing tools and differences in geographic, cultural, and age-related variance of social media use and communication. CASM utilizes a team-based approach to analysis of social media data, explicitly informed by community expertise. We use of CASM to analyze Twitter posts from gang-involved youth in Chicago. We designed a set of experiments to evaluate the performance of a support vector machine using CASM hand-labeled posts against a distant model. We found that the CASM-informed hand-labeled data outperforms the baseline distant labels, indicating that the CASM labels capture additional dimensions of information that content-only methods lack. We then question whether this is helpful or harmful for gun violence prevention.
Desmond Upton Patton, William R. Frey, Kyle A. McGregor, Fei-Tzin Lee, Kathy McKeown, Emanuel Moss
AIES5
2020 Event-Guided Denoising for Multilingual Relation Learning
abstract
General purpose relation extraction has recently seen considerable gains in part due to a massively data-intensive distant supervision technique from Soares et al. (2019)that produces stateof-the-art results across many benchmarks.In this work, we present a methodology for collecting high quality training data for relation extraction from unlabeled text that achieves a nearrecreation of their zero-shot and few-shot results at a fraction of the training cost.Our approach exploits the predictable distributional structure of date-marked news articles to build a denoised corpus -the extraction process filters out low quality examples.We show that a smaller multilingual encoder trained on this corpus performs comparably to the current state-of-the-art (when both receive little to no fine-tuning) on few-shot and standard relation benchmarks in English and Spanish despite using many fewer examples (50k vs. 300mil+).
Amith Ananthram, Emily Allaway, Kathy McKeown
COLING3
2020 Detecting Urgency Status of Crisis Tweets: A Transfer Learning Approach for Low Resource Languages
abstract
We release an urgency dataset that consists of English tweets relating to natural crises.The set is annotated along with annotations of their corresponding urgency status.Additionally, we release evaluation datasets for two low-resource languages, i.e.Sinhala and Odia, and demonstrate an effective zero-shot transfer from English to these two languages by training cross-lingual classifiers.We adopt cross-lingual embeddings constructed using different methods to extract features of the tweets, including a few state-of-the-art contextual embeddings such as BERT, RoBERTa and XLM-R.We train a variety of classifier architectures, supervised and semi supervised, on the extracted features.We also further experiment with ensembling the various classifiers.With very limited amounts of labeled data in English and zero data in the low resource languages, we show a successful framework of training monolingual and cross-lingual classifiers using deep learning methods which are known to be data hungry.Specifically, we show that the recent deep contextual embeddings are also helpful when dealing with very small-scale datasets.Classifiers that incorporate RoBERTa yield the best performance for the English urgency detection task, with 25% F1 score absolute improvement over the baselines.For the zero-shot transfer to low resource languages, classifiers that use LASER features perform the best for Sinhala transfer while XLM-R features benefit the Odia transfer the most.
Efsun Sarioglu Kayi, Linyong Nan, Bohan Qu, Mona T. Diab, Kathy McKeown
COLING5
2020 Zero-Shot Stance Detection: A Dataset and Model using Generalized Topic Representations
abstract
Stance detection is an important component of understanding hidden influences in everyday life.Since there are thousands of potential topics to take a stance on, most with little to no training data, we focus on zero-shot stance detection: classifying stance from no training examples.In this paper, we present a new dataset for zero-shot stance detection that captures a wider range of topics and lexical variation than in previous datasets.Additionally, we propose a new model for stance detection that implicitly captures relationships between topics using generalized topic representations and show that this model improves performance on a number of challenging linguistic phenomena.
Emily Allaway, Kathy McKeown
EMNLP (1)2
2020 Severing the Edge Between Before and After: Neural Architectures for Temporal Ordering of Events
abstract
Miguel Ballesteros, Rishita Anubhai, Shuai Wang, Nima Pourdamghani, Yogarshi Vyas, Jie Ma, Parminder Bhatia, Kathleen McKeown, Yaser Al-Onaizan. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Miguel Ballesteros, Rishita Anubhai, Nima Pourdamghani, Yogarshi Vyas, Jie Ma 0005, Parminder Bhatia, Kathy McKeown, Yaser Al-Onaizan
EMNLP (1)8
2020 Controllable Meaning Representation to Text Generation: Linearization and Data Augmentation Strategies
abstract
We study the degree to which neural sequenceto-sequence models exhibit fine-grained controllability when performing natural language generation from a meaning representation.Using two task-oriented dialogue generation benchmarks, we systematically compare the effect of four input linearization strategies on controllability and faithfulness.Additionally, we evaluate how a phrase-based data augmentation method can improve performance.We find that properly aligning input sequences during training leads to highly controllable generation, both when training from scratch or when fine-tuning a larger pre-trained model.Data augmentation further improves control on difficult, randomly generated utterance plans.
Chris Kedzie, Kathy McKeown
EMNLP (1)2
2019 Neural Network Alignment for Sentential Paraphrases
abstract
We present a monolingual alignment system for long, sentence-or clause-level alignments, and demonstrate that systems designed for word-or short phrase-based alignment are illsuited for these longer alignments.Our system is capable of aligning semantically similar spans of arbitrary length.We achieve significantly higher recall on aligning phrases of four or more words and outperform state-ofthe-art aligners on the long alignments in the MSR RTE corpus.
Jessica Ouyang 0001, Kathy McKeown
ACL (1)2
2019 AMPERSAND: Argument Mining for PERSuAsive oNline Discussions
abstract
Tuhin Chakrabarty, Christopher Hidey, Smaranda Muresan, Kathy McKeown, Alyssa Hwang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tuhin Chakrabarty, Christopher Hidey, Smaranda Muresan, Kathy McKeown, Alyssa Hwang
EMNLP/IJCNLP (1)4
2019 Automatically Inferring Gender Associations from Language
abstract
Serina Chang, Kathy McKeown. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Serina Chang, Kathy McKeown
EMNLP/IJCNLP (1)2
2019 Detecting and Reducing Bias in a High Stakes Domain
abstract
Ruiqi Zhong, Yanda Chen, Desmond Patton, Charlotte Selous, Kathy McKeown. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ruiqi Zhong, Yanda Chen, Desmond Upton Patton, Charlotte Selous, Kathy McKeown
EMNLP/IJCNLP (1)5
2019 Multimodal Social Media Analysis for Gang Violence Prevention
Philipp Blandfort, Desmond Upton Patton, William R. Frey, Svebor Karaman, Surabhi Bhargava, Fei-Tzin Lee, Siddharth Varia, Chris Kedzie, Michael B. Gaskell, Rossano Schifanella, Kathy McKeown, Shih-Fu Chang
ICWSM11
2019 A Good Sample is Hard to Find: Noise Injection Sampling and Self-Training for Neural Language Generation Models
abstract
Deep neural networks (DNN) are quickly becoming the de facto standard modeling method for many natural language generation (NLG) tasks.In order for such models to truly be useful, they must be capable of correctly generating utterances for novel meaning representations (MRs) at test time.In practice, even sophisticated DNNs with various forms of semantic control frequently fail to generate utterances faithful to the input MR.In this paper, we propose an architecture agnostic selftraining method to sample novel MR/text utterance pairs to augment the original training data.Remarkably, after training on the augmented data, even simple encoder-decoder models with greedy decoding are capable of generating semantically correct utterances that are as good as state-of-the-art outputs in both automatic and human evaluations of quality.
Chris Kedzie, Kathy McKeown
INLG2
2018 Persuasive Influence Detection: The Role of Argument Sequencing
abstract
Automatic detection of persuasion in online discussion is key to understanding how social media is used. Predicting persuasiveness is difficult, however, due to the need to model world knowledge, dialogue, and sequential reasoning. We focus on modeling the sequence of arguments in social media posts using neural models with embeddings for words, discourse relations, and semantic frames. We demonstrate significant improvement over prior work in detecting successful arguments. We also present an error analysis assessing novice human performance at predicting persuasiveness.
Christopher Hidey, Kathy McKeown
AAAI2
2018 Detecting Gang-Involved Escalation on Social Media Using Context
abstract
Serina Chang, Ruiqi Zhong, Ethan Adams, Fei-Tzin Lee, Siddharth Varia, Desmond Patton, William Frey, Chris Kedzie, Kathy McKeown. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018.
Serina Chang, Ruiqi Zhong, Ethan Adams, Fei-Tzin Lee, Siddharth Varia, Desmond Upton Patton, William R. Frey, Chris Kedzie, Kathy McKeown
EMNLP9
2018 Content Selection in Deep Learning Models of Summarization
abstract
We carry out experiments with deep learning models of summarization across the domains of news, personal stories, meetings, and medical articles in order to understand how content selection is performed.We find that many sophisticated features of state of the art extractive summarizers do not improve performance over simpler models.These results suggest that it is easier to create a summarizer for a new domain than previous work suggests and bring into question the benefit of deep learning models for summarization for those domains that do have massive datasets (i.e., news).At the same time, they suggest important questions for new research in summarization; namely, new forms of sentence representations or external knowledge sources are needed that are better suited to the sumarization task.
Chris Kedzie, Kathy McKeown, Hal Daumé III
EMNLP2
2018 Cross-lingual sentiment transfer with limited resources
Mohammad Sadegh Rasooli, Noura Farra, Axinia Radeva, Tao Yu 0009, Kathy McKeown
Mach. Transl.5
2017 SMARTies: Sentiment Models for Arabic Target entities
abstract
We consider entity-level sentiment analysis in Arabic, a morphologically rich language with increasing resources.We present a system that is applied to complex posts written in response to Arabic newspaper articles.Our goal is to identify important entity "targets" within the post along with the polarity expressed about each target.We achieve significant improvements over multiple baselines, demonstrating that the use of specific morphological representations improves the performance of identifying both important targets and their sentiment, and that the use of distributional semantic clusters further boosts performances for these representations, especially when richer linguistic resources are not available.
Noura Farra, Kathy McKeown
EACL (1)2
2017 Human-Centric Justification of Machine Learning Predictions
abstract
Human decision makers in many domains can make use of predictions made by machine learning models in their decision making process, but the usability of these predictions is limited if the human is unable to justify his or her trust in the prediction. We propose a novel approach to producing justifications that is geared towards users without machine learning expertise, focusing on domain knowledge and on human reasoning, and utilizing natural language generation. Through a task-based experiment, we show that our approach significantly helps humans to correctly decide whether or not predictions are accurate, and significantly increases their satisfaction with the justification.
Or Biran, Kathy McKeown
IJCAI2
2017 Domain-Adaptable Hybrid Generation of RDF Entity Descriptions
abstract
RDF ontologies provide structured data on entities in many domains and continue to grow in size and diversity. While they can be useful as a starting point for generating descriptions of entities, they often miss important information about an entity that cannot be captured as simple relations. In addition, generic approaches to generation from RDF cannot capture the unique style and content of specific domains. We describe a framework for hybrid generation of entity descriptions, which combines generation from RDF data with text extracted from a corpus, and extracts unique aspects of the domain from the corpus to create domain-specific generation systems. We show that each component of our approach significantly increases the satisfaction of readers with the text across multiple applications and domains.
Or Biran, Kathy McKeown
IJCNLP(1)2
2017 Detecting Influencers in Multiple Online Genres
abstract
Social media has become very popular and mainstream, leading to an abundance of content. This wealth of content contains many interactions and conversations that can be analyzed for a variety of information. One such type of information is analyzing the roles people take in a conversation. Detecting influencers, one such role, can be useful for political campaigning, successful advertisement strategies, and detecting terrorist leaders. We explore influence in discussion forums, weblogs, and micro-blogs through the development of learned language analysis components to recognize known indicators of influence. Our components are author traits, agreement, claims, argumentation, persuasion, credibility, and certain dialog patterns. Each of these components is motivated by social science through Robert Cialdini’s “Weapons of Influence” [Cialdini 2007]. We classify influencers across five online genres and analyze which features are most indicative of influencers in each genre. First, we describe a rich suite of features that were generated using each of the system components. Then, we describe our experiments and results, including using domain adaptation to exploit the data from multiple online genres.
Sara Rosenthal, Kathy McKeown
ACM Trans. Internet Techn.2
2016 Mining Paraphrasal Typed Templates from a Plain Text Corpus
abstract
Finding paraphrases in text is an important task with implications for generation, summarization and question answering, among other applications.Of particular interest to those applications is the specific formulation of the task where the paraphrases are templated, which provides an easy way to lexicalize one message in multiple ways by simply plugging in the relevant entities.Previous work has focused on mining paraphrases from parallel and comparable corpora, or mining very short sub-sentence synonyms and paraphrases.In this paper we present an approach which combines distributional and KB-driven methods to allow robust mining of sentence-level paraphrasal templates, utilizing a rich type system for the slots, from a plain text corpus.
Or Biran, Terra Blevins, Kathy McKeown
ACL (1)3
2016 Identifying Causal Relations Using Parallel Wikipedia Articles
Christopher Hidey, Kathy McKeown
ACL (1)2
2016 Automatically Processing Tweets from Gang-Involved Youth: Towards Detecting Loss and Aggression
abstract
Violence is a serious problems for cities like Chicago and has been exacerbated by the use of social media by gang-involved youths for taunting rival gangs. We present a corpus of tweets from a young and powerful female gang member and her communicators, which we have annotated with discourse intention, using a deep read to understand how and what triggered conversations to escalate into aggression. We use this corpus to develop a part-of-speech tagger and phrase table for the variant of English that is used and a classifier for identifying tweets that express grieving and aggression.
Terra Blevins, Robert Kwiatkowski, Jamie C. Macbeth, Kathy McKeown, Desmond Upton Patton, Owen Rambow
COLING4
2016 Real-Time Web Scale Event Summarization Using Sequential Decision Making
Chris Kedzie, Fernando Diaz 0001, Kathy McKeown
IJCAI3
2016 Extractive and Abstractive Event Summarization over Streaming Web Text
Chris Kedzie, Kathy McKeown
IJCAI2
2016 Predicting the impact of scientific concepts using full-text features
abstract
New scientific concepts, interpreted broadly, are continuously introduced in the literature, but relatively few concepts have a long‐term impact on society. The identification of such concepts is a challenging prediction task that would help multiple parties—including researchers and the general public—focus their attention within the vast scientific literature. In this paper we present a system that predicts the future impact of a scientific concept, represented as a technical term, based on the information available from recently published research articles. We analyze the usefulness of rich features derived from the full text of the articles through a variety of approaches, including rhetorical sentence analysis, information extraction, and time‐series analysis. The results from two large‐scale experiments with 3.8 million full‐text articles and 48 million metadata records support the conclusion that full‐text features are significantly more useful for prediction than metadata‐only features and that the most accurate predictions result from combining the metadata and full‐text features. Surprisingly, these results hold even when the metadata features are available for a much larger number of documents than are available for the full‐text features.
Kathy McKeown, Hal Daumé III, Snigdha Chaturvedi, John Paparrizos, Kapil Thadani, Pablo Barrio 0002, Or Biran, Suvarna Bothe, Michael Collins 0001, Kenneth R. Fleischmann, Luis Gravano, Rahul Jha, Ben King, Kevin McInerney, Taesun Moon, Arvind Neelakantan, Diarmuid Ó Séaghdha, Dragomir R. Radev, Thomas Clay Templeton, Simone Teufel
J. Assoc. Inf. Sci. Technol.1
2015 Predicting Salient Updates for Disaster Summarization
abstract
Chris Kedzie, Kathleen McKeown, Fernando Diaz. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Chris Kedzie, Kathy McKeown, Fernando Diaz 0001
ACL (1)2
2015 Discourse Planning with an N-gram Model of Relations
abstract
While it has been established that transitions between discourse relations are important for coherence, such information has not so far been used to aid in language generation.We introduce an approach to discourse planning for conceptto-text generation systems which simultaneously determines the order of messages and the discourse relations between them.This approach makes it straightforward to use statistical transition models, such as n-gram models of discourse relations learned from an annotated corpus.We show that using such a model significantly improves the quality of the generated text as judged by humans.
Or Biran, Kathy McKeown
EMNLP2
2015 System Combination for Machine Translation through Paraphrasing
abstract
In this paper, we propose a paraphrasing model to address the task of system combination for machine translation.We dynamically learn hierarchical paraphrases from target hypotheses and form a synchronous context-free grammar to guide a series of transformations of target hypotheses into fused translations.The model is able to exploit phrasal and structural system-weighted consensus and also to utilize existing information about word ordering present in the target hypotheses.In addition, to consider a diverse set of plausible fused translations, we develop a hybrid combination architecture, where we paraphrase every target hypothesis using different fusing techniques to obtain fused translations for each target, and then make the final selection among all fused translations.Our experimental results show that our approach can achieve a significant improvement over combination baselines.i h EP to denote i h E attached with related word positions, use i h e to denote a phrase within i h E , and use i h ep to denote i h e attached with related word positions.For instance, If i h E is "you buy the book", then i h EP would be "you 1 buy 2 the 3 book 4 ".If i h eis "the book", then i h ep is "the 3 book 4 ".For a given sentence i, a MT system h and a MT system k, we use a SCFG denoted by i
Wei-Yun Ma, Kathy McKeown
EMNLP2
2015 Modeling Reportable Events as Turning Points in Narrative
abstract
We present novel experiments in modeling the rise and fall of story characteristics within narrative, leading up to the Most Reportable Event (MRE), the compelling event that is the nucleus of the story.We construct a corpus of personal narratives from the bulletin board website Reddit, using the organization of Reddit content into topic-specific communities to automatically identify narratives.Leveraging the structure of Reddit comment threads, we automatically label a large dataset of narratives.We present a change-based model of narrative that tracks changes in formality, affect, and other characteristics over the course of a story, and we use this model in distant supervision and selftraining experiments that achieve significant improvements over the baselines at the task of identifying MREs.
Jessica Ouyang 0001, Kathy McKeown
EMNLP2
2015 PDTB Discourse Parsing as a Tagging Task: The Two Taggers Approach
abstract
Full discourse parsing in the PDTB framework is a task that has only recently been attempted.We present the Two Taggers approach, which reformulates the discourse parsing task as two simpler tagging tasks: identifying the relation within each sentence, and identifying the relation between each pair of adjacent sentences.We then describe a system that uses two CRFs to achieve an F1 score of 39.33, higher than the only previously existing system, at the full discourse parsing task.Our results show that sequential information is important for discourse relations, especially cross-sentence relations, and that a simple approach to argument span identification is enough to achieve state of the art results.We make our easy to use, easy to extend parser publicly available.
Or Biran, Kathy McKeown
SIGDIAL Conference2
2015 I Couldn't Agree More: The Role of Conversational Structure in Agreement and Disagreement Detection in Online Discussions
abstract
Determining when conversational participants agree or disagree is instrumental for broader conversational analysis; it is necessary, for example, in deciding when a group has reached consensus.In this paper, we describe three main contributions.We show how different aspects of conversational structure can be used to detect agreement and disagreement in discussion forums.In particular, we exploit information about meta-thread structure and accommodation between participants.Second, we demonstrate the impact of the features using 3-way classification, including sentences expressing disagreement, agreement or neither.Finally, we show how to use a naturally occurring data set with labels derived from the sides that participants choose in debates on createdebate.com.The resulting new agreement corpus, Agreement by Create Debaters (ABCD) is 25 times larger than any prior corpus.We demonstrate that using this data enables us to outperform the same system trained on prior existing in-domain smaller annotated datasets.
Sara Rosenthal, Kathy McKeown
SIGDIAL Conference2
2014 Towards Automatic Detection of Narrative Structure
Jessica Ouyang 0001, Kathy McKeown
LREC2
2013 Sentence Compression with Joint Structural Inference
Kapil Thadani, Kathy McKeown
CoNLL2
2013 Classifying Taxonomic Relations between Pairs of Wikipedia Articles
Or Biran, Kathy McKeown
IJCNLP2
2013 Cluster-based Web Summarization
Yves Petinot, Kathy McKeown, Kapil Thadani
IJCNLP2
2013 Supervised Sentence Fusion with Single-Stage Inference
Kapil Thadani, Kathy McKeown
IJCNLP2
2013 Using a Supertagged Dependency Language Model to Select a Good Translation in System Combination
Wei-Yun Ma, Kathy McKeown
HLT-NAACL2
2012 Can Automatic Post-Editing Make MT More Meaningful
Kristen Parton, Nizar Habash, Kathy McKeown, Gonzalo Iglesias, Adrià de Gispert
EAMT3
2012 Annotating Agreement and Disagreement in Threaded Discussion
Jacob Andreas, Sara Rosenthal, Kathy McKeown
LREC3
2011 Age Prediction in Blogs: A Study of Style, Content, and Online Behavior in Pre- and Post-Social Media Generations
Sara Rosenthal, Kathy McKeown
ACL2
2011 Identifying Event Descriptions using Co-training with Online News Summaries
William Yang Wang, Kapil Thadani, Kathy McKeown
IJCNLP3
2011 System Combination for Machine Translation Based on Text-to-Text Generation
Wei-Yun Ma, Kathy McKeown
MTSummit2
2011 Information Status Distinctions and Referring Expressions: An Empirical Study of References to People in News Summaries
abstract
Although there has been much theoretical work on using various information status distinctions to explain the form of references in written text, there have been few studies that attempt to automatically learn these distinctions for generating references in the context of computer-regenerated text. In this article, we present a model for generating references to people in news summaries that incorporates insights from both theory and a corpus analysis of human written summaries. In particular, our model captures how two properties of a person referred to in the summary—familiarity to the reader and global salience in the news story—affect the content and form of the initial reference to that person in a summary. We demonstrate that these two distinctions can be learned from a typical input for multi-document summarization and that they can be used to make regeneration decisions that improve the quality of extractive summaries.
Advaith Siddharthan, Ani Nenkova, Kathy McKeown
Comput. Linguistics3
2010 Automatic Attribution of Quoted Speech in Literary Narrative
abstract
We describe a method for identifying the speakers of quoted speech in natural-language textual stories. We have assembled a corpus of more than 3,000 quotations, whose speakers (if any) are manually identified, from a collection of 19th and 20th century literature by six authors. Using rule-based and statistical learning, our method identifies candidate characters, determines their genders, and attributes each quote to the most likely speaker. We divide the quotes into syntactic classes in order to leverage common discourse patterns, which enable rapid attribution for many quotes. We apply learning algorithms to the remainder and achieve an overall accuracy of 83%.
David K. Elson, Kathy McKeown
AAAI2
2010 Extracting Social Networks from Literary Fiction
David K. Elson, Nicholas Dames, Kathy McKeown
ACL3
2010 "Got You!": Automatic Vandalism Detection in Wikipedia with Web-based Shallow Syntactic-Semantic Modeling
William Yang Wang, Kathy McKeown
COLING2
2010 Tense and Aspect Assignment in Narrative Discourse
David K. Elson, Kathy McKeown
INLG2
2010 Building a Bank of Semantically Encoded Narratives
David K. Elson, Kathy McKeown
LREC2
2010 Towards Semi-Automated Annotation for Prepositional Phrase Attachment
Sara Rosenthal, William Lipovsky, Kathy McKeown, Kapil Thadani, Jacob Andreas
LREC3
2010 Time-Efficient Creation of an Accurate Sentence Fusion Corpus
Kathy McKeown, Sara Rosenthal, Kapil Thadani, Coleman Moore
HLT-NAACL1
2009 Who, What, When, Where, Why? Comparing Multiple Approaches to the Cross-Lingual 5W Task
Kristen Parton, Kathy McKeown, Bob Coyne, Mona T. Diab, Ralph Grishman, Dilek Hakkani-Tür, Mary P. Harper, Heng Ji 0001, Wei-Yun Ma, Adam Meyers 0001, Sara Stolbach, Ang Sun, Gökhan Tür, Wei Xu 0004, Sibel Yaman
ACL/IJCNLP2
2009 Contextual Phrase-Level Polarity Analysis Using Lexical Affect Scoring and Syntactic N-Grams
Apoorv Agarwal, Fadi Biadsy, Kathy McKeown
EACL3
2009 Classification-based strategies for combining multiple 5-w question answering systems
abstract
We describe and analyze inference strategies for combining outputs from multiple question answering systems each of which was developed independently. Specifically, we address the DARPA-funded GALE information distillation Year 3 task of finding answers to the 5-Wh questions (who, what, when, where, and why) for each given sentence. The approach we take revolves around determining the best system using discriminative learning. In particular, we train support vector machines with a set of novel features that encode systems’ capabilities of returning as many correct answers as possible. We analyze two combination strategies: one combines multiple systems at the granularity of sentences, and the other at the granularity of individual fields. Our experimental results indicate that the proposed features and combination strategies were able to improve the overall performance by 22% to 36% relative to a random selection, 16% to 35% relative to a majority voting scheme, and 15% to 23% relative to the best individual system. Index Terms: Question answering, Systems for spoken language understanding
Sibel Yaman, Dilek Hakkani-Tür, Gökhan Tür, Ralph Grishman, Mary P. Harper, Kathy McKeown, Adam Meyers 0001, Kartavya Sharma
INTERSPEECH6
2008 Simultaneous multilingual search for translingual information retrieval
abstract
We consider the problem of translingual information retrieval, where monolingual searchers issue queries in a different language than the document language(s) and the results must be returned in the language they know, the query language. We present a framework for translingual IR that integrates document translation and query translation into the retrieval model. The corpus is represented as an aligned, jointly indexed "pseudo-parallel" corpus, where each document contains the text of the document along with its translation into the query language. The queries are formulated as multilingual structured queries, where each query term and its translations into the document language(s) are treated as synonym sets. This model leverages simultaneous search in multiple languages against jointly indexed documents to improve the accuracy of results over search using document translation or query translation alone. For query translation, we compared a statistical machine translation (SMT) approach to a dictionary-based approach. We found that using a Wikipedia-derived dictionary for named entities combined with an SMT-based dictionary worked better than SMT alone. Simultaneous multilingual search also has other important features suited to translingual search, since it can provide an indication of poor document translation when a match with the source document is found. We show how close integration of CLIR and SMT allows us to improve result translation in addition to IR results.
Kristen Parton, Kathy McKeown, James Allan 0001, Enrique Henestroza
CIKM2
2008 A Framework for Identifying Textual Redundancy
Kapil Thadani, Kathy McKeown
COLING2
2007 Using Question-Answer Pairs in Extractive Summarization of Email Conversations
Kathy McKeown, Lokesh Shrestha, Owen Rambow
CICLing1
2007 Building and Refining Rhetorical-Semantic Relation Models
Sasha Blair-Goldensohn, Kathy McKeown, Owen Rambow
HLT-NAACL2
2007 Lexicalized Markov Grammars for Sentence Compression
Michel Galley, Kathy McKeown
HLT-NAACL2
2007 Question Answering Using Integrated Information Retrieval and Information Extraction
Barry Schiffman, Kathy McKeown, Ralph Grishman, James Allan 0001
HLT-NAACL2
2006 Automatic Creation of Domain Templates
Elena Filatova, Vasileios Hatzivassiloglou, Kathy McKeown
ACL3
2006 Lessons Learned from Large Scale Evaluation of Systems that Produce Text: Nightmares and Pleasant Surprises
Kathy McKeown
INLG1
2006 A compositional context sensitive multi-document summarizer: exploring the factors that influence summarization
abstract
The usual approach for automatic summarization is sentence extraction, where key sentences from the input documents are selected based on a suite of features. While word frequency often is used as a feature in summarization, its impact on system performance has not been isolated. In this paper, we study the contribution to summarization of three factors related to frequency: content word frequency, composition functions for estimating sentence importance from word frequency, and adjustment of frequency weights based on context. We carry out our analysis using datasets from the Document Understanding Conferences, studying not only the impact of these features on automatic summarizers, but also their role in human summarization. Our research shows that a frequency based summarizer can achieve performance comparable to that of state-of-the-art systems, but only with a good composition function; context sensitivity improves performance and significantly reduces repetition.
Ani Nenkova, Lucy Vanderwende, Kathy McKeown
SIGIR3
2005 Facilitating Physicians' Access to Information via Tailored Text Summarization
Noémie Elhadad, Kathy McKeown, David R. Kaufman, Desmond A. Jordan
AMIA2
2005 From text to speech summarization
abstract
In this paper, we present approaches used in text summarization, showing how they can be adapted for speech summarization and where they fall short. Informal style and apparent lack of structure in speech mean that the typical approaches used for text summarization must be extended for use with speech. We illustrate how features derived from speech can help determine summary content within two ongoing summarization projects at Columbia University.
Kathy McKeown, Julia Hirschberg, Michel Galley, Sameer Maskey
ICASSP (5)1
2005 Do summaries help?
abstract
We describe a task-based evaluation to determine whether multi-document summaries measurably improve user performance whe using online news browsing systems for directed research. We evaluated the multi-document summaries generated by Newsblaster, a robust news browsing system that clusters online news articles and summarizes multiple articles on each event. Four groups of subjects were asked to perform the same time-restricted fact-gathering tasks, reading news under different conditions: no summaries at all, single sentence summaries drawn from one of the articles, Newsblaster multi-document summaries, and human summaries. Our results show that, in comparison to source documents only, the quality of reports assembled using Newsblaster summaries was significantly better and user satisfaction was higher with both Newsblaster and human summaries.
Kathy McKeown, Rebecca J. Passonneau, David K. Elson, Ani Nenkova, Julia Hirschberg
SIGIR1
2005 Customization in a unified framework for summarizing medical literature
Noémie Elhadad, Min-Yen Kan, Judith L. Klavans, Kathy McKeown
Artif. Intell. Medicine4
2005 Sentence Fusion for Multidocument News Summarization
abstract
A system that can produce informative summaries, highlighting common information found in many online documents, will help Web users to pinpoint information that they need without extensive reading. In this article, we introduce sentence fusion, a novel text-to-text generation technique for synthesizing common information across documents. Sentence fusion involves bottom-up local multisequence alignment to identify phrases conveying similar information and statistical generation to combine common phrases into a sentence. Sentence fusion moves the summarization field from the use of purely extractive methods to the generation of abstracts that contain sentences not found in any of the input documents and can synthesize information across sources.
Regina Barzilay, Kathy McKeown
Comput. Linguistics2
2004 Identifying Agreement and Disagreement in Conversational Speech: Use of Bayesian Networks to Model Pragmatic Dependencies
abstract
We describe a statistical approach for modeling agreements and disagreements in conversational interaction. Our approach first identifies adjacency pairs using maximum entropy ranking based on a set of lexical, durational, and structural features that look both forward and backward in the discourse. We then classify utterances as agreement or disagreement using these adjacency pairs and features that represent various pragmatic influences of previous agreement or disagreement on the current utterance. Our approach achieves 86.9% accuracy, a 4.9% increase over previous work.
Michel Galley, Kathy McKeown, Julia Hirschberg, Elizabeth Shriberg
ACL2
2004 Detection of Question-Answer Pairs in Email Conversations
Lokesh Shrestha, Kathy McKeown
COLING2
2004 Syntactic Simplification for Improving Content Selection in Multi-Document Summarization
Advaith Siddharthan, Ani Nenkova, Kathy McKeown
COLING3
2004 Generating Overview Summaries of Ongoing Email Thread Discussions
Stephen Wan 0001, Kathy McKeown
COLING2
2003 Discourse Segmentation of Multi-Party Conversation
abstract
We present a domain-independent topic segmentation algorithm for multi-party speech. Our feature-based algorithm combines knowledge about content using a text-based algorithm as a feature and about form using linguistic and acoustic cues about topic shifts extracted from speech. This segmentation algorithm uses automatically induced decision rules to combine the different features. The embedded text-based algorithm builds on lexical cohesion and has performance comparable to state-of-the-art algorithms based on lexical information. A significant error reduction is obtained by combining the two knowledge sources.
Michel Galley, Kathy McKeown, Eric Fosler-Lussier, Hongyan Jing
ACL2
2003 Statistical Acquisition of Content Selection Rules for Natural Language Generation
Pablo Ariel Duboue, Kathy McKeown
EMNLP2
2003 Improving Word Sense Disambiguation in Lexical Chaining
Michel Galley, Kathy McKeown
IJCAI2
2003 PROGENIE: Biographical Descriptions for Intelligence Analysis
Pablo Ariel Duboue, Kathy McKeown, Vasileios Hatzivassiloglou
ISI2
2003 Columbia's Newsblaster: New Features and Future Directions
Kathy McKeown, Regina Barzilay, David K. Elson, David Kirk Evans, Judith L. Klavans, Ani Nenkova, Barry Schiffman, Sergey Sigelman
HLT-NAACL1
2003 References to Named Entities: a Corpus Study
Ani Nenkova, Kathy McKeown
HLT-NAACL2
2003 DefScriber: a hybrid system for definitional QA
abstract
No abstract available.
Sasha Blair-Goldensohn, Kathy McKeown, Andrew Hazen Schlaikjer
SIGIR2
2002 Usability evaluation of an experimental text summarization system and three search engines: implications for the reengineering of health care interfaces
Andre Kushniruk, Min-Yen Kan, Kathy McKeown, Judith L. Klavans, Desmond A. Jordan, Mark Laflamme, Vimla L. Patel
AMIA3
2002 NLP Found Helpful (at least for one Text Categorization Task)
abstract
Attempts to use natural language processing (NLP) for text categorization and information retrieval (IR) have had mixed results. Nevertheless, there is a strong intuition that NLP is important, at least for some tasks. In this paper, we discuss a task involving captioned images for which the subject and the predicate are critical. The usefulness of NLP for this task is established in two ways. In addition to the standard method of introducing a new system and comparing its performance with others in the literature, we also present evidence from experiments with human subjects showing that NLP generally improves speed and accuracy.
Carl L. Sable, Kathy McKeown, Kenneth Church 0001
EMNLP2
2002 Content Planner Construction via Evolutionary Algorithms and a Corpus-based Fitness Function
Pablo Ariel Duboue, Kathy McKeown
INLG2
2002 Corpus-trained Text Generation for Summarization
Min-Yen Kan, Kathy McKeown
INLG2
2002 Using the Annotated Bibliography as a Resource for Indicative Summarization
Min-Yen Kan, Judith L. Klavans, Kathy McKeown
LREC3
2002 Introduction to the Special Issue on Summarization
abstract
generation based on rhetorical structure extraction. In Proceedings of the International Conference on Computational Linguistics, Kyoto, Japan, pages 344–348. Otterbacher, Jahna, Dragomir R. Radev, and Airong Luo. 2002. Revisions that improve cohesion in multi-document summaries: A preliminary study. In ACL Workshop on Text Summarization, Philadelphia. Papineni, K., S. Roukos, T. Ward, and W-J. Zhu. 2001. BLEU: A method for automatic evaluation of machine translation. Research Report RC22176, IBM. Radev, Dragomir, Simone Teufel, Horacio Saggion, Wai Lam, John Blitzer, Arda Celebi, Hong Qi, Elliott Drabek, and Danyu Liu. 2002. Evaluation of text summarization in a cross-lingual information retrieval framework. Technical Report, Center for Language and Speech Processing, Johns Hopkins University, Baltimore, June. Radev, Dragomir R., Hongyan Jing, and Malgorzata Budzikowska. 2000. Centroid-based summarization of multiple documents: Sentence extraction, utility-based evaluation, and user studies. In ANLP/NAACL Workshop on Summarization, Seattle, April. Radev, Dragomir R. and Kathleen R. McKeown. 1998. Generating natural language summaries from multiple on-line sources. Computational Linguistics, 24(3):469–500. Rau, Lisa and Paul Jacobs. 1991. Creating segmented databases from free text for text retrieval. In Proceedings of the 14th Annual International ACM-SIGIR Conference on Research and Development in Information Retrieval, New York, pages 337–346. Saggion, Horacio and Guy Lapalme. 2002. Generating indicative-informative summaries with SumUM. Computational Linguistics, 28(4), 497–526. Salton, G., A. Singhal, M. Mitra, and C. Buckley. 1997. Automatic text structuring and summarization. Information Processing & Management, 33(2):193–207. Silber, H. Gregory and Kathleen McCoy. 2002. Efficiently computed lexical chains as an intermediate representation for automatic text summarization. Computational Linguistics, 28(4), 487–496. Sparck Jones, Karen. 1999. Automatic summarizing: Factors and directions. In I. Mani and M. T. Maybury, editors, Advances in Automatic Text Summarization. MIT Press, Cambridge, pages 1–13. Strzalkowski, Tomek, Gees Stein, J. Wang, and Bowden Wise. 1999. A robust practical text summarizer. In I. Mani and M. T. Maybury, editors, Advances in Automatic Text Summarization. MIT Press, Cambridge, pages 137–154. Teufel, Simone and Marc Moens. 2002. Summarizing scientific articles: Experiments with relevance and rhetorical status. Computational Linguistics, 28(4), 409–445. White, Michael and Claire Cardie. 2002. Selecting sentences for multidocument summaries using randomized local search. In Proceedings of the Workshop on Automatic Summarization (including DUC 2002), Philadelphia, July. Association for Computational Linguistics, New Brunswick, NJ, pages 9–18. Witbrock, Michael and Vibhu Mittal. 1999. Ultra-summarization: A statistical approach to generating highly condensed non-extractive summaries. In Proceedings of the 22nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Berkeley, pages 315–316. Zechner, Klaus. 2002. Automatic summarization of open-domain multiparty dialogues in diverse genres. Computational Linguistics, 28(4), 447–485.
Dragomir R. Radev, Eduard H. Hovy, Kathy McKeown
Comput. Linguistics3
2002 Exploring features from natural language generation for prosody modeling
Shimei Pan, Kathy McKeown, Julia Hirschberg
Comput. Speech Lang.2
2002 Inferring Strategies for Sentence Ordering in Multidocument News Summarization
abstract
The problem of organizing information for multidocument summarization so that the generated summary is coherent has received relatively little attention. While sentence ordering for single document summarization can be determined from the ordering of sentences in the input article, this is not the case for multidocument summarization where summary sentences may be drawn from different input articles. In this paper, we propose a methodology for studying the properties of ordering information in the news genre and describe experiments done on a corpus of multiple acceptable orderings we developed for the task. Based on these experiments, we implemented a strategy for ordering information that combines constraints from chronological order of events and topical relatedness. Evaluation of our augmented algorithm shows a significant improvement of the ordering over two baseline strategies.
Regina Barzilay, Noémie Elhadad, Kathy McKeown
J. Artif. Intell. Res.3
2001 Extracting Paraphrases from a Parallel Corpus
abstract
While paraphrasing is critical both for interpretation and generation of natural language, current systems use manual or semi-automatic methods to collect paraphrases. We present an unsupervised learning algorithm for identification of paraphrases from a corpus of multiple English translations of the same source text. Our approach yields phrasal and single word lexical paraphrases as well as syntactic paraphrases.
Regina Barzilay, Kathy McKeown
ACL2
2001 Empirically Estimating Order Constraints for Content Planning in Generation
abstract
In a language generation system, a content planner embodies one or more "plans" that are usually hand-crafted, sometimes through manual analysis of target text.In this paper, we present a system that we developed to automatically learn elements of a plan and the ordering constraints among them.As training data, we use semantically annotated transcripts of domain experts performing the task our system is designed to mimic.Given the large degree of variation in the spoken language of the transcripts, we developed a novel algorithm to find parallels between transcripts based on techniques used in computational genomics.Our proposed methodology was evaluated two-fold: the learning and generalization capabilities were quantitatively evaluated using cross validation obtaining a level of accuracy of 89%.A qualitative evaluation is also provided.
Pablo Ariel Duboue, Kathy McKeown
ACL2
2001 Re-engineering an Inference Engine to Support Continuous Quality Improvement
Pablo Ariel Duboue, Desmond A. Jordan, Kathy McKeown
AMIA3
2001 Personalizing retrieval of journal articles for patient care
Simone Teufel, Vasileios Hatzivassiloglou, Kathy McKeown, Desmond A. Jordan, Kathleen M. Dunn, Sergey Sigelman, Andre Kushniruk
AMIA3
2001 Semantic abnormality and its realization in spoken language
abstract
In this paper we investigate the relationship between various lexical and prosodic features and semantic abnormality, the occurrence of unusual or unexpected events, in generating speech for MAGIC, which employs a Concept-to-Speech system to generate post-operative reports for patients who have undergone bypass surgery. Using the speech corpus collected for this application, we conducted empirical analysis to systematically discover significantly correlated prosodic and lexical features. The automatically learned abnormality model not only can be used in building comprehensive prosody prediction systems for Concept-to-Speech generation, but also help identify unusual information during speech analysis and understanding.
Shimei Pan, Kathy McKeown, Julia Hirschberg
INTERSPEECH2
2001 Research Paper: Generation and Evaluation of Intraoperative Inferences for Automated Health Care Briefings on Patient Status After Bypass Surgery
abstract
OBJECTIVE: The authors present a system that scans electronic records from cardiac surgery and uses inference rules to identify and classify abnormal events (e.g., hypertension) that may occur during critical surgical points (e.g., start of bypass). This vital information is used as the content of automatically generated briefings designed by MAGIC, a multimedia system that they are developing to brief intensive care unit clinicians on patient status after cardiac surgery. By recognizing patterns in the patient record, inferences concisely summarize detailed patient data. DESIGN: The authors present the development of inference rules that identify important information about patient status and describe their implementation and an experiment they carried out to validate their correctness. The data for a set of 24 patients were analyzed independently by the system and by 46 physicians. MEASUREMENTS: The authors measured accuracy, specificity, and sensitivity by comparing system inferences against physician judgments, in cases where all three physicians agreed and against the majority opinion in all cases. RESULTS: For laboratory inferences, evaluation shows that the system has an average accuracy of 98 percent (full agreement) and 96 percent (majority model). An analysis of interrater agreement, however, showed that physicians do not agree on abnormal hemodynamic events and could not serve as a gold standard for evaluating hemodynamic events. Analysis of discrepancies reveals possibilities for system improvement and causes of physician disagreement. CONCLUSIONS: This evaluation shows that the laboratory inferences of the system have high accuracy. The lack of agreement among physicians highlights the need for an objective quality-assurance tool for hemodynamic inferences. The system provides such a tool by implementing inferencing procedures established in the literature.
Desmond A. Jordan, Kathy McKeown, Kristian J. Concepcion, Steven K. Feiner, Vasileios Hatzivassiloglou
J. Am. Medical Informatics Assoc.2
2000 A study of communication in the Cardiac Surgery Intensive Care Unit and its implications for automated briefing
Kathy McKeown, Desmond A. Jordan, Steven K. Feiner, James Shaw, Elizabeth S. Chen, Shabina Ahmad, Andre Kushniruk, Vimla L. Patel
AMIA1
2000 Experiments in Automated Lexicon Building for Text Searching
Barry Schiffman, Kathy McKeown
COLING2
2000 Integrating a Large-Scale, Reusable Lexicon with a Natural Language Generator
abstract
This paper presents the integration of a large-scale, reusable lexicon for generation with the FUF/SURGE unification-based syntactic realizer. The lexicon was combined from multiple existing resources in a semi-automatic process. The integration is a multi-step unification process. This integration allows the reuse of lexical, syntactic, and semantic knowledge encoded in the lexicon in the development of lexical chooser module in a generation system. The lexicon also brings other benefits to a generation system: for example, the ability to generate many lexical and syntactic paraphrases and the ability to avoid non-grammatical output.
Hongyan Jing, Yael Dahan Netzer, Michael Elhadad, Kathy McKeown
INLG4
2000 Generating Referring Quantified Expressions
abstract
In this paper, we describe how quantifiers can be generated in a text generation system. By taking advantage of discourse and ontological information, quantified expressions can replace entities in a text, making the text more fluent and concise. In addition to avoiding ambiguities between distributive and collective readings in universal quantification generation, we will also show how different scope orderings between universal and existential quantifiers will result in different quantified expressions in our algorithm.
James Shaw, Kathy McKeown
INLG2
2000 Learning Methods to Combine Linguistic Indicators: Improving Aspectual Classification and Revealing Linguistic Insights
abstract
Aspectual classification maps verbs to a small set of primitive categories in order to reason about time. This classification is necessary for interpreting temporal modifiers and assessing temporal relationships, and is therefore a required component for many natural language applications. A verb's aspectual category can be predicted by co-occurrence frequencies between the verb and certain linguistic modifiers. These frequency measures, called linguistic indicators, are chosen by linguistic insights. However, linguistic indicators used in isolation are predictively incomplete, and are therefore insufficient when used individually. In this article, we compare three supervised machine learning methods for combining multiple linguistic indicators for aspectual classification: decision trees, genetic programming, and logistic regression. A set of 14 indicators are combined for classification according to two aspectual distinctions. This approach improves the classification performance for both distinctions, as evaluated over unrestricted sets of verbs occurring across two corpora. This demonstrates the effectiveness of the linguistic indicators and provides a much-needed full-scale method for automatic aspectual classification. Moreover, the models resulting from learning reveal several linguistic insights that are relevant to aspectual classification. We also compare supervised learning methods with an unsupervised method for this task.
Eric V. Siegel, Kathy McKeown
Comput. Linguistics2
1999 Information Fusion in the Context of Multi-Document Summarization
abstract
We present a method to automatically generate a concise summary by identifying and synthesizing similar elements across related text from a set of multiple documents. Our approach is unique in its usage of language generation to reformulate the wording of the summary.
Regina Barzilay, Kathy McKeown, Michael Elhadad
ACL2
1999 Extracting Patient Profiles from Patient Records and Online Literature
Vasileios Hatzivassiloglou, Olga Merport, Kathy McKeown, Desmond A. Jordan
AMIA3
1999 Word Informativeness and Automatic Pitch Accent Modeling
Shimei Pan, Kathy McKeown
EMNLP2
1999 The Decomposition of Human-Written Summary Sentences
abstract
We define the problem of decomposing human-written summary sentences and propose a novel Hidden Markov Model solution to the problem.Human summarizers often rely on cutting and pasting of the full document to generate summaries.Decomposing a human-written summary sentence requires determining: (1) whether it is constructed by cutting and pasting, (2) what components in the sentence come from the original document, and (3) where in the document the components come from.Solving the decomposition problem can potentially lead to the automatic acquisition of large corpora for summarization.It also sheds light on the generation of summary text by cutting and pasting.The evaluation shows that the proposed decomposition algorithm performs well.
Hongyan Jing, Kathy McKeown
SIGIR2
1998 Generating Multimedia Briefings: Coordinating Language and Illustration
abstract
Communication can be more effective when several media (such as text, speech, or graphics) are integrated and coordinated to present information. This changes the nature of media-specific generation (e.g., language or graphics generation), which must take into account the multimedia context in which it occurs. This paper presents work on coordinating and integrating speech, text, static and animated three-dimensional graphics, and stored images, as part of several systems we have developed at Columbia University. A particular focus of our work has been on the generation of presentations that brief a user on information of interest
Kathy McKeown, Steven K. Feiner, Mukesh Dalal, Shih-Fu Chang
Artif. Intell.1
1998 Generating Natural Language Summaries from Multiple On-Line Sources
Dragomir R. Radev, Kathy McKeown
Comput. Linguistics2
1997 Predicting the Semantic Orientation of Adjectives
abstract
We identify and validate from a large corpus constraints from conjunctions on the positive or negative semantic orientation of the conjoined adjectives. A log-linear regression model uses these constraints to predict whether conjoined adjectives are of same or different orientations, achieving 82% accuracy in this task when each conjunction is considered independently. Combining the constraints across many adjectives, a clustering algorithm separates the adjectives into groups of different orientations, and finally, adjectives are labeled positive or negative. Evaluations on real data and simulation experiments indicate high levels of performance: classification precision is more than 90% for adjectives that occur in a modest number of conjunctions in the corpus.
Vasileios Hatzivassiloglou, Kathy McKeown
ACL2
1997 Generating Multimedia Briefings: Language Generation in a Coordinated Multimedia Environment
Kathy McKeown
IJCAI1
1997 Floating Constraints in Lexical Choice
Michael Elhadad, Kathy McKeown, Jacques Robin
Comput. Linguistics2
1997 A Technical Word- and Term-Translation Aid Using Noisy Parallel Corpora across Language Groups
Pascale Fung, Kathy McKeown
Mach. Transl.2
1996 Spoken language generation in a multimedia system
Shimei Pan, Kathy McKeown
ICSLP2
1996 Negotiation for Automated Generation of Temporal Multimedia Presentations
abstract
Creating high-quality multimedia presentations requires much skill, time, and effort.This is particularly true when temporal media, such as speech and animation, are involved.We describe the design and implementation of a knowledge-based system that generates customized temporal multimedia presentations.We provide art overview of the system's architecture, and explain how speech, written text, and graphics are generated and coordinated.Our emphasis is on how temporal media are coordinated by the system through a multi-stage negotiation process.In negotiation, media-specific generation components interact with a novel coordination component that solves temporal constraints provided by the generators.We illustrate our work with a set of examples generated by the system in a testbed application intended to update hospital caregivers on the status of patients who have undergone a cardiac bypass operation.
Mukesh Dalal, Steven K. Feiner, Kathy McKeown, Shimei Pan, Michelle X. Zhou, Tobias Höllerer, James Shaw, Jeanne C. Fromer
ACM Multimedia3
1996 Empirically Designing and Evaluating a New Revision-Based Model for Summary Generation
abstract
We present a system for summarizing quantitative data in natural language, focusing on the use of a corpus of basketball game summaries, drawn from on-line news services, to empirically shape the system design and to evaluate our approach. Our initial corpus analysis revealed characteristics of textual summaries that challenge the capabilities of current language generation systems. In order to meet these challenges, we developed a revision-based model for summary generation and implemented it in our prototype system streak. A second, detailed corpus analysis was used to identify and encode the revision rules of the system. Finally, we carried out a quantitative evaluation, using several test corpora, to measure the robustness of the new revision-based model. Our results show that our new model improves both coverage and extensibility of the traditional language generation model.
Jacques Robin, Kathy McKeown
Artif. Intell.2
1996 Translating Collocations for Bilingual Lexicons: A Statistical Approach
Frank Smadja, Kathy McKeown, Vasileios Hatzivassiloglou
Comput. Linguistics2
1995 A Quantitative Evaluation of Linguistic Tests for the Automatic Prediction of Semantic Markedness
abstract
We present a corpus-based study of methods that have been proposed in the linguistics literature for selecting the semantically unmarked term out of a pair of antonymous adjectives. Solutions to this problem are applicable to the more general task of selecting the positive term from the pair. Using automatically collected data, the accuracy and applicability of each method is quantified, and a statistical analysis of the significance of the results is performed. We show that some simple methods are indeed good indicators for the answer to the problem while other proposed methods fail to perform better than would be attributable to chance. In addition, one of the simplest methods, text frequency, dominates all others. We also apply two generic statistical learning methods for combining the indications of the individual methods, and compare their performance to the simple methods. The most sophisticated complex learning method offers a small, but statistically significant, improvement over the original tests.
Vasileios Hatzivassiloglou, Kathy McKeown
ACL2
1995 Generating Summaries of Multiple News Articles
abstract
We present a natural language system which summarizes a series of news articles on the same event. It uses summarization operators, identified through empirical analysis of a corpus of news summaries, to group empirical analysis of a corpus of news summaries, to group together templates from the output of the systems developed for ARPA's Message Understanding Conferences. Depending on the available resources(e.g. space), summaries of different length can be produced. Our research also provides a methodological frame work for future work on the summarization task and on the evaluation of new summarization systems.
Kathy McKeown, Dragomir R. Radev
SIGIR1
1995 Generating Concise Natural Language Summaries
abstract
Summaries typically convey maximal information in minimal space. In this paper, we describe an approach to summary generation that opportunistically folds information from multiple facts into a single sentence using concise linguistic constructions. Unlike previous work in generation, how information gets added into a summary depends in part on constraints from how the text is worded so far. This approach allows the construction of concise summaries, containing complex sentences that pack in information. The resulting summary sentences are, in fact, longer than sentences generated by previous systems. We describe two applications we have developed using this approach, one of which produces summaries of basketball games (STREAK) while the other (PLANDOC) produces summaries of telephone network planning activity; both systems summarize input data as opposed to full text. The applications implement opportunistic summary generation using complementary approaches. STREAK uses revision, creating a draft of essential facts and then using revision rules constrained by the draft wording to add in additional facts as the text allows. PLANDOC uses discourse planning, looking ahead in its text plan to group together facts which can be expressed concisely using conjunction and deleting repetitions. In this paper, we describe the problems for summary generation, the two domains, the linguistic constructions that the systems use to convey information concisely and the textual constraints that determine what information gets included.
Kathy McKeown, Jacques Robin, Karen Kukich
Inf. Process. Manag.1
1995 The challenge of spoken language systems: research directions for the nineties
abstract
A spoken language system combines speech recognition, natural language processing and human interface technology. It functions by recognizing the person's words, interpreting the sequence of words to obtain a meaning in terms of the application, and providing an appropriate response back to the user. Potential applications of spoken language systems range from simple tasks, such as retrieving information from an existing database (traffic reports, airline schedules), to interactive problem solving tasks involving complex planning and reasoning (travel planning, traffic routing), to support for multilingual interactions. We examine eight key areas in which basic research is needed to produce spoken language systems: (1) robust speech recognition; (2) automatic training and adaptation; (3) spontaneous speech; (4) dialogue models; (5) natural language response generation; (6) speech synthesis and speech generation; (7) multilingual systems; and (8) interactive multimodal systems. In each area, we identify key research challenges, the infrastructure needed to support research, and the expected benefits. We conclude by reviewing the need for multidisciplinary research, for development of shared corpora and related resources, for computational support and far rapid communication among researchers. The successful development of this technology will increase accessibility of computers to a wide range of users, will facilitate multinational communication and trade, and will create new research specialties and jobs in this rapidly expanding area.>
Ronald A. Cole, Lynette Hirschman, Les E. Atlas, Mary E. Beckman, Alan Biermann, Marcia A. Bush, Mark A. Clements, Jordan Cohen, Oscar Garcia, Brian A. Hanson, Hynek Hermansky, Steve Levinson, Kathy McKeown, Nelson Morgan, David G. Novick, Mari Ostendorf, Sharon L. Oviatt, Patti Price, Harvey F. Silverman, Judy Spitz, Alex Waibel, Clifford J. Weinstein, Stephen A. Zahorian, Victor Zue
IEEE Trans. Speech Audio Process.13
1994 Emergent Linguistic Rules from inducing Decision Trees: Disambiguating Discourse Clue Words
Eric V. Siegel, Kathy McKeown
AAAI2
1993 Corpus Analysis for Revision-Based Generation of Complex Sentences
Jacques Robin, Kathy McKeown
AAAI2
1993 Towards the Automatic Identification of Adjectival Scales: Clustering Adjectives According to Meaning
abstract
In this paper we present a method to group adjectives according to their meaning, as a first step towards the automatic identification of adjectival scales. We discuss the properties of adjectival scales and of groups of semantically related adjectives and how they imply sources of linguistic knowledge in text corpora. We describe how our system exploits this linguistic knowledge to compute a measure of similarity between two adjectives, using statistical techniques and without having access to any semantic information about the adjectives. We also show how a clustering algorithm can use these similarities to produce the groups of adjectives, and we present results produced by our system for a sample set of adjectives. We conclude by presenting evaluation methods for the task at hand, and analyzing the significance of the results obtained.
Vasileios Hatzivassiloglou, Kathy McKeown
ACL2
1993 Tailoring Lexical Choice to the User's Vocabulary in Multimedia Explanation Generation
abstract
In this paper, we discuss the different strategies used in COMET (COordinated Multimedia Explanation Testbed) for selecting words with which the user is familiar. When pictures cannot be used to disambiguate a word or phrase, COMET has four strategies for avoiding unknown words. We give examples for each of these strategies and show how they are implemented in COMET.
Kathy McKeown, Jacques Robin, Michael A. Tanenblatt
ACL1
1992 Generating Cross-References for Multimedia Explanation
Kathy McKeown, Steven K. Feiner, Jacques Robin, Dorée D. Seligmann, Michael A. Tanenblatt
AAAI1
1991 COMET: generating coordinated multimedia explanations
abstract
No abstract available.
Steven K. Feiner, Kathy McKeown
CHI2
1990 Coordinating Text and Graphics in Explanation Generation
Steven K. Feiner, Kathy McKeown
AAAI2
1990 User Models and User Interfaces
Kathy McKeown
AAAI1
1990 Automatically Extracting and Representing Collocations for Language Generation
abstract
Collocational knowledge is necessary for language generation. The problem is that collocations come in a large variety of forms. They can involve two, three or more words, these words can be of different syntactic categories and they can be involved in more or less rigid ways. This leads to two main difficulties: collocational knowledge has to be acquired and it must be represented flexibly so that it can be used for language generation. We address both problems in this paper, focusing on the acquisition problem. We describe a program, Xtract, that automatically acquires a range of collocations from large textual corpora and we describe how they can be represented in a flexible lexicon using a unification based formalism.
Frank Smadja, Kathy McKeown
ACL2
1990 Generating Connectives
Michael Elhadad, Kathy McKeown
COLING2
1988 Beyond Semantic Ambiguity
Galina Datskovsky Moerdler, Kathy McKeown
AAAI2
1987 Functional Unification Grammar Revisited
abstract
In this paper, we show that one benefit of FUG, the ability to state global constraints on choice separately from syntactic rules, is difficult in generation systems based on augmented context free grammars (e.g., Definite Clause Grammars). They require that such constraints be expressed locally as part of syntactic rules and therefore, duplicated in the grammar. Finally, we discuss a reimplementation of FUG that achieves the similar levels of efficiency as Rubinoff's adaptation of MUMBLE, a deterministic language generator.
Kathy McKeown, Cécile Paris
ACL1
1987 Building Natural Language Interfaces for Rule-based Expert Systems
Galina Datskovsky Moerdler, Kathy McKeown, J. Robert Ensor
IJCAI2
1986 Language generation: Applications, issues, and approaches
abstract
As computer systems become more sophisticated they must be able to communicate their results successfully to their users. Natural language generation is the area of research concerned with developing methods that will allow a computer system to respond to its user in human language. In this paper, the need for natural language generation is first motivated by showing how it is used in several applications. Given that language generation is necessary for such systems, the paper also focuses on the issues that must be taken into account in developing a system that can generate language. Finally, techniques that have been used in two question-answering systems, the TEXT system [21] and TAILOR [22], are discussed.
Kathy McKeown
Proc. IEEE1
1985 Tailoring Explanations for the User
Kathy McKeown, Myron Wish, Kevin Matthews
IJCAI1
1985 Discourse Strategies for Generating Natural-Language Text
Kathy McKeown
Artif. Intell.1
1984 Using Focus to Generate Complex and Simple Sentences
abstract
One problem for the generation of natural language text is determining when to use a sequence of simple sentences and when a single complex one is more appropriate. In this paper, we show how focus of attention is one factor that influences this decision and describe its implementation in a system that generates explanations for a student advisor expert system. The implementation uses tests on functional information such as focus of attention within the Prolog definite clause grammar formalism to determine when to use complex sentences, resulting in an efficient generator that has the same benefits as a functional grammar system.
Marcia A. Derr, Kathy McKeown
COLING2
1984 Natural Language for Exert Systems: Comparisons with Database Systems
abstract
Article Free Access Share on Natural language for expert systems: comparisons with database systems Author: Kathleen R. McKeown Columbia University, New York, N.Y. Columbia University, New York, N.Y.View Profile Authors Info & Claims ACL '84/COLING '84: Proceedings of the 10th International Conference on Computational Linguistics and 22nd annual meeting on Association for Computational LinguisticsJuly 1984 Pages 190–193https://doi.org/10.3115/980491.980534Published:02 July 1984Publication History 2citation335DownloadsMetricsTotal Citations2Total Downloads335Last 12 Months16Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Kathy McKeown
COLING1
1983 Recursion in TEXT and Its Use in Language Generation
Kathy McKeown
AAAI1
1983 Focus Constraints on Language Generation
Kathy McKeown
IJCAI1
1983 Paraphrasing Questions Using Given and New Information
Kathy McKeown
Am. J. Comput. Linguistics1
1982 The Text System for Natural Language Generation: an Overview
abstract
Computer-based generation of natural language requires consideration of two different types of problems: 1) determining the content and textual shape of what is to be said, and 2) transforming that message into English. A computational solution to the problems of deciding what to say and how to organize it effectively is proposed that relies on an interaction between structural and semantic processes. Schemas, which encode aspects of discourse structure, are used to guide the generation process. A focusing mechanism monitors the use of the schemas, providing constraints on what can be said at any point. These mechanisms have been implemented as part of a generation method within the context of a natural language database system, addressing the specific problem of responding to questions about database structure.
Kathy McKeown
ACL1
1980 Generating Relevant Explanations: Natural Language Responses to Questions about Database Structure
Kathy McKeown
AAAI1
1980 Creating polyhedral stellations
abstract
A process for creating and displaying stellations of a given polyhedral solid is described. A stellation is one of many star-like polyhedra which can be derived from a single solid by extending its existing faces. A program has been implemented which performs the stellation process on an input object and generates a 3-dimensional image of the stellated object on a computer graphics display screen. Pictures of icosahedron and rhombictriacontahedron stellations generated by the program are included in the paper.
Kathy McKeown, Norman I. Badler
SIGGRAPH1
1979 Paraphrasing Using Given and New Information in a Question-Answer System
abstract
The design and implementation of a paraphrase component for a natural language question-answer system (CO-OP) is presented. A major point made is the role of given and new information in formulating a paraphrase that differs in a meaningful way from the user's question. A description is also given of the transformational grammar used by the paraphraser to generate questions.
Kathy McKeown
ACL1