Ani Nenkova

dblp:58/896 · DBLP profile ↗
← Back
82ranked-venue papers
11as first author
16since 2021 · last 2024
0000-0002-5825-7875ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 71 · 10 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2024 Few-Shot Dialogue Summarization via Skeleton-Assisted Prompt Transfer in Prompt Tuning
abstract
Kaige Xie, Tong Yu, Haoliang Wang, Junda Wu, Handong Zhao, Ruiyi Zhang, Kanak Mahadik, Ani Nenkova, Mark Riedl. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Kaige Xie, Tong Yu 0001, Junda Wu, Handong Zhao, Ruiyi Zhang 0002, Kanak Mahadik, Ani Nenkova, Mark O. Riedl
EACL (1)8
2024 SOHES: Self-supervised Open-world Hierarchical Entity Segmentation
abstract
Open-world entity segmentation, as an emerging computer vision task, aims at segmenting entities in images without being restricted by pre-defined classes, offering impressive generalization capabilities on unseen images and concepts. Despite its promise, existing entity segmentation methods like Segment Anything Model (SAM) rely heavily on costly expert annotators. This work presents Self-supervised Open-world Hierarchical Entity Segmentation (SOHES), a novel approach that eliminates the need for human annotations. SOHES operates in three phases: self-exploration, self-instruction, and self-correction. Given a pre-trained self-supervised representation, we produce abundant high-quality pseudo-labels through visual feature clustering. Then, we train a segmentation model on the pseudo-labels, and rectify the noises in pseudo-labels via a teacher-student mutual-learning procedure. Beyond segmenting entities, SOHES also captures their constituent parts, providing a hierarchical understanding of visual entities. Using raw images as the sole training data, our method achieves unprecedented performance in self-supervised open-world segmentation, marking a significant milestone towards high-quality open-world entity segmentation in the absence of human-annotated masks. Project page: https://SOHES.github.io.
Shengcao Cao, Jiuxiang Gu, Jason Kuen, Hao Tan 0002, Ruiyi Zhang 0002, Handong Zhao, Ani Nenkova, Liangyan Gui, Tong Sun 0005, Yu-Xiong Wang
ICLR7
2024 ADOPD: A Large-Scale Document Page Decomposition Dataset
abstract
Research in document image understanding is hindered by limited high-quality document data. To address this, we introduce ADOPD, a comprehensive dataset for document page decomposition. ADOPD stands out with its data-driven approach for document taxonomy discovery during data collection, complemented by dense annotations. Our approach integrates large-scale pretrained models with a human-in-the-loop process to guarantee diversity and balance in the resulting data collection. Leveraging our data-driven document taxonomy, we collect and densely annotate document images, addressing four document image understanding tasks: Doc2Mask, Doc2Box, Doc2Tag, and Doc2Seq. Specifically, for each image, the annotations include human-labeled entity masks, text bounding boxes, as well as automatically generated tags and captions that have been manually cleaned. We conduct comprehensive experimental analyses to validate our data and assess the four tasks using various models. We envision ADOPD as a foundational dataset with the potential to drive future research in document understanding.
Jiuxiang Gu, Xiangxi Shi, Jason Kuen, Ruiyi Zhang 0002, Anqi Liu 0001, Ani Nenkova, Tong Sun 0005
ICLR7
2023 Factual or Contextual? Disentangling Error Types in Entity Description Generation
abstract
In the task of entity description generation, given a context and a specified entity, a model must describe that entity correctly and in a contextually-relevant way.In this task, as well as broader language generation tasks, the generation of a nonfactual description (factual error) versus an incongruous description (contextual error) is fundamentally different, yet often conflated.We develop an evaluation paradigm that enables us to disentangle these two types of errors in naturally occurring textual contexts.We find that factuality and congruity are often at odds, and that models specifically struggle with accurate descriptions of entities that are less familiar to people.This shortcoming of language models raises concerns around the trustworthiness of such models, since factual errors on less well-known entities are exactly those that a human reader will not recognize.1 1 The code and data used in the paper is available at https: //github.com/navitagoyal/Factual-or-Contextual-E rrors-in-LM-Desc-Gen.
Navita Goyal, Ani Nenkova, Hal Daumé III
ACL (1)2
2023 Learning the Visualness of Text Using Large Vision-Language Models
abstract
Visual text evokes an image in a person's mind, while non-visual text fails to do so.A method to automatically detect visualness in text will enable text-to-image retrieval and generation models to augment text with relevant images.This is particularly challenging with long-form text as text-to-image generation and retrieval models are often triggered for text that is designed to be explicitly visual in nature, whereas long-form text could contain many non-visual sentences.To this end, we curate a dataset of 3,620 English sentences and their visualness scores provided by multiple human annotators.We also propose a fine-tuning strategy that adapts large vision-language models like CLIP by modifying the model's contrastive learning objective to map text identified as nonvisual to a common NULL image while matching visual text to their corresponding images in the document.We evaluate the proposed approach on its ability to (i) classify visual and non-visual text accurately, and (ii) attend over words that are identified as visual in psycholinguistic studies.Empirical evaluation indicates that our approach performs better than several heuristics and baseline models for the proposed task.Furthermore, to highlight the importance of modeling the visualness of text, we conduct qualitative analyses of text-to-image generation systems like DALL-E.
Gaurav Verma 0005, Ryan Rossi, Chris Tensmeyer, Jiuxiang Gu, Ani Nenkova
EMNLP5
2023 Summaries as Captions: Generating Figure Captions for Scientific Documents with Automated Text Summarization
abstract
Chieh-Yang Huang, Ting-Yao Hsu, Ryan Rossi, Ani Nenkova, Sungchul Kim, Gromit Yeuk-Yin Chan, Eunyee Koh, C Lee Giles, Ting-Hao Huang. Proceedings of the 16th International Natural Language Generation Conference. 2023.
Chieh-Yang Huang, Ting-Yao Hsu, Ryan Rossi, Ani Nenkova, Sungchul Kim, Gromit Yeuk-Yin Chan, Eunyee Koh, C. Lee Giles, Ting-Hao 'Kenneth' Huang
INLG4
2023 LayerDoc: Layer-wise Extraction of Spatial Hierarchical Structure in Visually-Rich Documents
abstract
Digital documents often contain images and scanned text. Parsing such visually-rich documents is a core task for work-flow automation, but it remains challenging since most documents do not encode explicit layout information, e.g., how characters and words are grouped into boxes and ordered into larger semantic entities. Current state-of-the-art layout extraction methods are challenged by such documents as they rely on word sequences to have correct reading order and do not exploit their hierarchical structure. We propose LayerDoc, an approach that uses visual features, textual semantics, and spatial coordinates along with constraint inference to extract the hierarchical layout structure of documents in a bottom-up layer-wise fashion. LayerDoc recursively groups smaller regions into larger semantic elements in 2D to infer complex nested hierarchies. Experiments show that our approach outperforms competitive baselines by 10-15% on three diverse datasets of forms and mobile app screen layouts for the tasks of spatial region classification, higher-order group identification, layout hierarchy extraction, reading order detection, and word grouping.
Puneet Mathur, Rajiv Jain, Ashutosh Mehra 0002, Jiuxiang Gu, Franck Dernoncourt, Anandhavelu Natarajan, Quan Hung Tran, Verena Kaynig, Ani Nenkova, Dinesh Manocha, Vlad I. Morariu
WACV9
2023 Web Table Formatting Affects Readability on Mobile Devices
abstract
Reading large tables on small mobile screens presents serious usability challenges that can be addressed, in part, by better table formatting. However, there are few evidenced-based guidelines for formatting mobile tables to improve readability. For this work, we first conducted a survey to investigate how people interact with tables on mobile devices and conducted a study with designers to identify which design considerations are most critical. Based on these findings, we designed and conducted three large scale studies with remote crowdworker participants. Across the studies, we analyze over 14,000 trials from 590 participants who each viewed and answered questions about 28 diverse tables rendered in different formats. We find that smaller cell padding and frozen headers lead to faster task completion, and that while zebra striping and row borders do not speed up tasks, they are still subjectively preferred by participants.
Chris Tensmeyer, Zoya Bylinskii, Tianyuan Cai 0004, David Bryan Miller, Ani Nenkova, Aleena Gertrudes Niklaus, Shaun Wallace
WWW5
2022 MGDoc: Pre-training with Multi-granular Hierarchy for Document Image Understanding
abstract
Zilong Wang, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios, Ani Nenkova, Tong Sun, Jingbo Shang, Vlad Morariu. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Zilong Wang 0002, Jiuxiang Gu, Chris Tensmeyer, Nikolaos Barmpalios, Ani Nenkova, Tong Sun 0005, Jingbo Shang, Vlad I. Morariu
EMNLP5
2022 DocLayoutTTS: Dataset and Baselines for Layout-informed Document-level Neural Speech Synthesis
Puneet Mathur, Franck Dernoncourt, Quan Hung Tran, Jiuxiang Gu, Ani Nenkova, Vlad I. Morariu, Rajiv Jain, Dinesh Manocha
INTERSPEECH5
2022 DI-2022: The Third Document Intelligence Workshop
abstract
Business documents are central to the operation of all organizations, and they come in all shapes and sizes: project reports, planning documents, technical specifications, financial statements, meeting minutes, legal agreements, contracts, resumes, purchase orders, invoices, and many more. The ability to read, understand and interpret these documents, referred to here as Document Intelligence (DI), is challenging due to not only many domains of knowledge involved, but also their complex formats and structures, internal and external cross references deployed, and even less-than-ideal quality of scans and OCR oftentimes performed on them. This workshop aims to explore and advance the current state of research and practice in answering these challenges.
Ani Nenkova, Douglas Burdick, Benjamin Han, Dave Lewis 0003, Sandeep Tata, Dan Tecuci
KDD1
2022 DocTime: A Document-level Temporal Dependency Graph Parser
abstract
Puneet Mathur, Vlad Morariu, Verena Kaynig-Fittkau, Jiuxiang Gu, Franck Dernoncourt, Quan Tran, Ani Nenkova, Dinesh Manocha, Rajiv Jain. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Puneet Mathur, Vlad I. Morariu, Verena Kaynig, Jiuxiang Gu, Franck Dernoncourt, Quan Hung Tran, Ani Nenkova, Dinesh Manocha, Rajiv Jain
NAACL-HLT7
2022 Temporal Effects on Pre-trained Models for Language Processing Tasks
abstract
Abstract Keeping the performance of language technologies optimal as time passes is of great practical interest. We study temporal effects on model performance on downstream language tasks, establishing a nuanced terminology for such discussion and identifying factors essential to conduct a robust study. We present experiments for several tasks in English where the label correctness is not dependent on time and demonstrate the importance of distinguishing between temporal model deterioration and temporal domain adaptation for systems using pre-trained representations. We find that, depending on the task, temporal model deterioration is not necessarily a concern. Temporal domain adaptation, however, is beneficial in all cases, with better performance for a given time period possible when the system is trained on temporally more recent data. Therefore, we also examine the efficacy of two approaches for temporal domain adaptation without human annotations on new data. Self-labeling shows consistent improvement and notably, for named entity recognition, leads to better temporal adaptation than even human annotations.
Oshin Agarwal, Ani Nenkova
Trans. Assoc. Comput. Linguistics2
2021 From Toxicity in Online Comments to Incivility in American News: Proceed with Caution
abstract
The ability to quantify incivility online, in news and in congressional debates, is of great interest to political scientists.Computational tools for detecting online incivility for English are now fairly accessible and potentially could be applied more broadly.We test the Jigsaw Perspective API for its ability to detect the degree of incivility on a corpus that we developed, consisting of manual annotations of civility in American news.We demonstrate that toxicity models, as exemplified by Perspective, are inadequate for the analysis of incivility in news.We carry out error analysis that points to the need to develop methods to remove spurious correlations between words often mentioned in the news, especially identity descriptors and incivility.Without such improvements, applying Perspective or similar models on news is likely to lead to wrong conclusions, that are not aligned with the human perception of incivility.
Anushree Hede, Oshin Agarwal, Linda Lu, Diana C. Mutz, Ani Nenkova
EACL5
2021 UniDoc: Unified Pretraining Framework for Document Understanding
abstract
Document intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabeled document datasets have opened up promising directions towards reducing annotation efforts by training models with self-supervised objectives. However, most of the existing document pretraining methods are still language-dominated. We present UDoc, a new unified pretraining framework for document understanding. UDoc is designed to support most document understanding tasks, extending the Transformer to take multimodal embeddings as input. Each input element is composed of words and visual features from a semantic region of the input document image. An important feature of UDoc is that it learns a generic representation by making use of three self-supervised losses, encouraging the representation to model sentences, learn similarities, and align modalities. Extensive empirical analysis demonstrates that the pretraining procedure learns better joint representations and leads to improvements in downstream tasks.
Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Nikolaos Barmpalios, Ani Nenkova, Tong Sun 0005
NeurIPS7
2021 Interpretability Analysis for Named Entity Recognition to Understand System Predictions and How They Can Improve
abstract
Abstract Named entity recognition systems achieve remarkable performance on domains such as English news. It is natural to ask: What are these models actually learning to achieve this? Are they merely memorizing the names themselves? Or are they capable of interpreting the text and inferring the correct entity type from the linguistic context? We examine these questions by contrasting the performance of several variants of architectures for named entity recognition, with some provided only representations of the context as features. We experiment with GloVe-based BiLSTM-CRF as well as BERT. We find that context does influence predictions, but the main factor driving high performance is learning the named tokens themselves. Furthermore, we find that BERT is not always better at recognizing predictive contexts compared to a BiLSTM-CRF model. We enlist human annotators to evaluate the feasibility of inferring entity types from context alone and find that humans are also mostly unable to infer entity types for the majority of examples on which the context-only system made errors. However, there is room for improvement: A system should be able to recognize any named entity in a predictive context correctly and our experiments indicate that current systems may be improved by such capability. Our human study also revealed that systems and humans do not always learn the same contextual clues, and context-only systems are sometimes correct even when humans fail to recognize the entity type from the context. Finally, we find that one issue contributing to model errors is the use of “entangled” representations that encode both contextual and local token information into a single vector, which can obscure clues. Our results suggest that designing models that explicitly operate over representations of local inputs and context, respectively, may in some cases improve performance. In light of these and related findings, we highlight directions for future work.
Oshin Agarwal, Yinfei Yang, Byron C. Wallace, Ani Nenkova
Comput. Linguistics4
2020 Trialstreamer: A living, automatically updated database of clinical trial reports
abstract
OBJECTIVE: Randomized controlled trials (RCTs) are the gold standard method for evaluating whether a treatment works in health care but can be difficult to find and make use of. We describe the development and evaluation of a system to automatically find and categorize all new RCT reports. MATERIALS AND METHODS: Trialstreamer continuously monitors PubMed and the World Health Organization International Clinical Trials Registry Platform, looking for new RCTs in humans using a validated classifier. We combine machine learning and rule-based methods to extract information from the RCT abstracts, including free-text descriptions of trial PICO (populations, interventions/comparators, and outcomes) elements and map these snippets to normalized MeSH (Medical Subject Headings) vocabulary terms. We additionally identify sample sizes, predict the risk of bias, and extract text conveying key findings. We store all extracted data in a database, which we make freely available for download, and via a search portal, which allows users to enter structured clinical queries. Results are ranked automatically to prioritize larger and higher-quality studies. RESULTS: As of early June 2020, we have indexed 673 191 publications of RCTs, of which 22 363 were published in the first 5 months of 2020 (142 per day). We additionally include 304 111 trial registrations from the International Clinical Trials Registry Platform. The median trial sample size was 66. CONCLUSIONS: We present an automated system for finding and categorizing RCTs. This yields a novel resource: a database of structured information automatically extracted for all published RCTs in humans. We make daily updates of this database available on our website (https://trialstreamer.robotreviewer.net).
Iain James Marshall, Benjamin E. Nye, Joël Kuiper, Anna Noel-Storr, Rachel Marshall, Rory Maclean, Frank Soboczenski, Ani Nenkova, James Thomas 0001, Byron C. Wallace
J. Am. Medical Informatics Assoc.8
2019 The Feasibility of Embedding Based Automatic Evaluation for Single Document Summarization
abstract
Simeng Sun, Ani Nenkova. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Simeng Sun, Ani Nenkova
EMNLP/IJCNLP (1)2
2018 A Corpus with Multi-Level Annotations of Patients, Interventions and Outcomes to Support Language Processing for Medical Literature
abstract
We present a corpus of 5,000 richly annotated abstracts of medical articles describing clinical randomized controlled trials. Annotations include demarcations of text spans that describe the Patient population enrolled, the Interventions studied and to what they were Compared, and the Outcomes measured (the 'PICO' elements). These spans are further annotated at a more granular level, e.g., individual interventions within them are marked and mapped onto a structured medical vocabulary. We acquired annotations from a diverse set of workers with varying levels of expertise and cost. We describe our data collection process and the corpus itself in detail. We then outline a set of challenging NLP tasks that would aid searching of the medical literature and the practice of evidence-based medicine.
Benjamin E. Nye, Junyi Jessy Li, Roma Patel, Yinfei Yang, Iain James Marshall, Ani Nenkova, Byron C. Wallace
ACL (1)6
2018 Evaluating Multiple System Summary Lengths: A Case Study
abstract
Practical summarization systems are expected to produce summaries of varying lengths, per user needs.While a couple of early summarization benchmarks tested systems across multiple summary lengths, this practice was mostly abandoned due to the assumed cost of producing reference summaries of multiple lengths.In this paper, we raise the research question of whether reference summaries of a single length can be used to reliably evaluate system summaries of multiple lengths.For that, we have analyzed a couple of datasets as a case study, using several variants of the ROUGE metric that are standard in summarization evaluation.Our findings indicate that the evaluation protocol in question is indeed competitive.This result paves the way to practically evaluating varying-length summaries with simple, possibly existing, summarization benchmarks.
Ori Shapira, David Gabay, Hadar Ronen, Judit Bar-Ilan, Yael Amsterdamer, Ani Nenkova, Ido Dagan
EMNLP6
2017 Aggregating and Predicting Sequence Labels from Crowd Annotations
abstract
Despite sequences being core to NLP, scant work has considered how to handle noisy sequence labels from multiple annotators for the same text. Given such annotations, we consider two complementary tasks: (1) aggregating sequential crowd labels to infer a best single set of consensus annotations; and (2) using crowd annotations as training data for a model that can predict sequences in unannotated text. For aggregation, we propose a novel Hidden Markov Model variant. To predict sequences in unannotated text, we propose a neural approach using Long Short Term Memory. We evaluate a suite of methods across two different applications and text genres: Named-Entity Recognition in news articles and Information Extraction from biomedical abstracts. Results show improvement over strong baselines. Our source code and data are available online.
An T. Nguyen 0001, Byron C. Wallace, Junyi Jessy Li, Ani Nenkova, Matthew Lease
ACL (1)4
2017 Combining Lexical and Syntactic Features for Detecting Content-Dense Texts in News
abstract
Content-dense news report important factual information about an event in direct, succinct manner. Information seeking applications such as information extraction, question answering and summarization normally assume all text they deal with is content-dense. Here we empirically test this assumption on news articles from the business, U.S. international relations, sports and science journalism domains. Our findings clearly indicate that about half of the news texts in our study are in fact not content-dense and motivate the development of a supervised content-density detector. We heuristically label a large training corpus for the task and train a two-layer classifying model based on lexical and unlexicalized syntactic features. On manually annotated data, we compare the performance of domain-specific classifiers, trained on data only from a given news domain and a general classifier in which data from all four domains is pooled together. Our annotation and prediction experiments demonstrate that the concept of content density varies depending on the domain and that naive annotators provide judgement biased toward the stereotypical domain label. Domain-specific classifiers are more accurate for domains in which content-dense texts are typically fewer. Domain independent classifiers reproduce better naive crowdsourced judgements. Classification prediction is high across all conditions, around 80%.
Yinfei Yang, Ani Nenkova
J. Artif. Intell. Res.2
2016 Improving the Annotation of Sentence Specificity
Junyi Jessy Li, Bridget O'Daniel, Wenli Zhao, Ani Nenkova
LREC5
2016 The Instantiation Discourse Relation: A Corpus Analysis of Its Properties and Improved Detection
abstract
INSTANTIATION is a fairly common discourse relation and past work has suggested that it plays special roles in local coherence, in sentiment expression and in content selection in summarization.In this paper we provide the first systematic corpus analysis of the relation and show that relation-specific features can improve considerably the detection of the relation.We show that sentences involved in INSTANTIATION are set apart from other sentences by the use of gradable (subjective) adjectives, the occurrence of rare words and by different patterns in part-of-speech usage.Words across arguments of INSTANTI-ATION are connected through hypernym and meronym relations significantly more often than in other sentences and that they stand out in context by being significantly less similar to each other than other adjacent sentence pairs.These factors provide substantial predictive power that improves the identification of implicit INSTANTIATION relation by more than 5% F-measure.
Junyi Jessy Li, Ani Nenkova
HLT-NAACL2
2015 Fast and Accurate Prediction of Sentence Specificity
abstract
Recent studies have demonstrated that specificity is an important characterization of texts potentially beneficial for a range of applications such as multi-document news summarization and analysis of science journalism. The feasibility of automatically predicting sentence specificity from a rich set of features has also been confirmed in prior work. In this paper we present a practical system for predicting sentence specificity which exploits only features that require minimum processing and is trained in a semi-supervised manner. Our system outperforms the state-of-the-art method for predicting sentence specificity and does not require part of speech tagging or syntactic parsing as the prior methods did. With the tool that we developed --- Speciteller --- we study the role of specificity in sentence simplification. We show that specificity is a useful indicator for finding sentences that need to be simplified and a useful objective for simplification, descriptive of the differences between original and simplified sentences.
Junyi Jessy Li, Ani Nenkova
AAAI2
2015 System Combination for Multi-document Summarization
abstract
We present a novel framework of system combination for multi-document summarization.For each input set (input), we generate candidate summaries by combining whole sentences from the summaries generated by different systems.We show that the oracle among these candidates is much better than the summaries that we have combined.We then present a supervised model to select among the candidates.The model relies on a rich set of features that capture content importance from different perspectives.Our model performs better than the systems that we combined based on manual and automatic evaluations.We also achieve very competitive performance on six DUC/TAC datasets, comparable to the state-of-the-art on most datasets.
Kai Hong, Mitchell P. Marcus, Ani Nenkova
EMNLP3
2015 Detecting Content-Heavy Sentences: A Cross-Language Case Study
abstract
The information conveyed by some sentences would be more easily understood by a reader if it were expressed in multiple sentences.We call such sentences content heavy: these are possibly grammatical but difficult to comprehend, cumbersome sentences.In this paper we introduce the task of detecting content-heavy sentences in cross-lingual context.Specifically we develop methods to identify sentences in Chinese for which English speakers would prefer translations consisting of more than one sentence.We base our analysis and definitions on evidence from multiple human translations and reader preferences on flow and understandability.We show that machine translation quality when translating content heavy sentences is markedly worse than overall quality and that this type of sentence are fairly common in Chinese news.We demonstrate that sentence length and punctuation usage in Chinese are not sufficient clues for accurately detecting heavy sentences and present a richer classification model that accurately identifies these sentences.
Junyi Jessy Li, Ani Nenkova
EMNLP2
2015 Identification and Characterization of Newsworthy Verbs in World News
abstract
We present a data-driven technique for acquiring domain-level importance of verbs from the analysis of abstract/article pairs of world news articles.We show that existing lexical resources capture some the semantic characteristics for important words in the domain.We develop a novel characterization of the association between verbs and personal story narratives, which is descriptive of verbs avoided in summaries for this domain.
Benjamin E. Nye, Ani Nenkova
HLT-NAACL2
2015 Inducing Lexical Style Properties for Paraphrase and Genre Differentiation
abstract
We present an intuitive and effective method for inducing style scores on words and phrases. We exploit signal in a phrase’s rate of occurrence across stylistically contrasting corpora, making our method simple to implement and efficient to scale. We show strong results both intrinsically, by correlation with human judgements, and extrinsically, in applications to genre analysis and paraphrasing.
Ellie Pavlick, Ani Nenkova
HLT-NAACL2
2015 Acoustic and lexical representations for affect prediction in spontaneous conversations
Houwei Cao, Arman Savran, Ragini Verma, Ani Nenkova
Comput. Speech Lang.4
2015 Speaker-sensitive emotion recognition via ranking: Studies on acted and spontaneous speech
Houwei Cao, Ragini Verma, Ani Nenkova
Comput. Speech Lang.3
2015 Temporal Bayesian Fusion for Affect Sensing: Combining Video, Audio, and Lexical Modalities
abstract
The affective state of people changes in the course of conversations and these changes are expressed externally in a variety of channels, including facial expressions, voice, and spoken words. Recent advances in automatic sensing of affect, through cues in individual modalities, have been remarkable; yet emotion recognition is far from a solved problem. Recently, researchers have turned their attention to the problem of multimodal affect sensing in the hope that combining different information sources would provide great improvements. However, reported results fall short of the expectations, indicating only modest benefits and occasionally even degradation in performance. We develop temporal Bayesian fusion for continuous real-value estimation of valence, arousal, power, and expectancy dimensions of affect by combining video, audio, and lexical modalities. Our approach provides substantial gains in recognition performance compared to previous work. This is achieved by the use of a powerful temporal prediction model as prior in Bayesian fusion as well as by incorporating uncertainties about the unimodal predictions. The temporal prediction model makes use of time correlations on the affect sequences and employs estimated temporal biases to control the affect estimations at the beginning of conversations. In contrast to other recent methods for combination of modalities our model is simpler, since it does not model relationships between modalities and involves only a few interpretable parameters to be estimated from the training data.
Arman Savran, Houwei Cao, Ani Nenkova, Ragini Verma
IEEE Trans. Cybern.3
2014 Detecting Information-Dense Texts in Multiple News Domains
abstract
We introduce the task of identifying information-dense texts,which report important factual information in direct, succinct manner. We describe a procedure that allows us to label automatically a large training corpus of New York Times texts.We train a classifier based on lexical, discourse and unlexicalized syntactic features and test its performance on a set of manually annotated articles from business, U.S. international relations, sports and science domains. Our results indicate that the task is feasible and that both syntactic and lexicalfeatures are highly predictive for the distinction. We observe considerable variation of prediction accuracy across domains and find that domain-specific models are more accurate.
Yinfei Yang, Ani Nenkova
AAAI2
2014 Cross-lingual Discourse Relation Analysis: A corpus study and a semi-supervised classification system
Junyi Jessy Li, Marine Carpuat, Ani Nenkova
COLING3
2014 Improving the Estimation of Word Importance for News Multi-Document Summarization
abstract
We introduce a supervised model for predicting word importance that incorporates a rich set of features.Our model is superior to prior approaches for identifying words used in human summaries.Moreover we show that an extractive summarizer using these estimates of word importance is comparable in automatic evaluation with the state-of-the-art.
Kai Hong, Ani Nenkova
EACL2
2014 Verbose, Laconic or Just Right: A Simple Computational Model of Content Appropriateness under Length Constraints
abstract
Length constraints impose implicit requirements on the type of content that can be included in a text.Here we propose the first model to computationally assess if a text deviates from these requirements.Specifically, our model predicts the appropriate length for texts based on content types present in a snippet of constant length.We consider a range of features to approximate content type, including syntactic phrasing, constituent compression probability, presence of named entities, sentence specificity and intersentence continuity.Weights for these features are learned using a corpus of summaries written by experts and on high quality journalistic writing.During test time, the difference between actual and predicted length allows us to quantify text verbosity.We use data from manual evaluation of summarization systems to assess the verbosity scores produced by our model.We show that the automatic verbosity scores are significantly negatively correlated with manual content quality scores given to the summaries.
Annie Louis, Ani Nenkova
EACL2
2014 A Repository of State of the Art and Competitive Baseline Summaries for Generic News Summarization
Kai Hong, John M. Conroy, Benoît Favre, Alex Kulesza, Ani Nenkova
LREC6
2014 Addressing Class Imbalance for Improved Recognition of Implicit Discourse Relations
abstract
In this paper we address the problem of skewed class distribution in implicit dis-course relation recognition. We examine the performance of classifiers for both bi-nary classification predicting if a particu-lar relation holds or not and for multi-class prediction. We review prior work to point out that the problem has been addressed differently for the binary and multi-class problems. We demonstrate that adopting a unified approach can significantly im-prove the performance of multi-class pre-diction. We also propose an approach that makes better use of the full annotations in the training set when downsampling is used. We report significant absolute im-provements in performance in multi-class prediction, as well as significant improve-ment of binary classifiers for detecting the presence of implicit Temporal, Compari-son and Contingency relations. 1
Junyi Jessy Li, Ani Nenkova
SIGDIAL Conference2
2014 Reducing Sparsity Improves the Recognition of Implicit Discourse Relations
abstract
The earliest work on automatic detec-tion of implicit discourse relations relied on lexical features. More recently, re-searchers have demonstrated that syntactic features are superior to lexical features for the task. In this paper we re-examine the two classes of state of the art representa-tions: syntactic production rules and word pair features. In particular, we focus on the need to reduce sparsity in instance repre-sentation, demonstrating that different rep-resentation choices even for the same class of features may exacerbate sparsity issues and reduce performance. We present re-sults that clearly reveal that lexicalization of the syntactic features is necessary for good performance. We introduce a novel, less sparse, syntactic representation which leads to improvement in discourse rela-tion recognition. Finally, we demonstrate that classifiers trained on different repre-sentations, especially lexical ones, behave rather differently and thus could likely be combined in future systems. 1
Junyi Jessy Li, Ani Nenkova
SIGDIAL Conference2
2014 CREMA-D: Crowd-Sourced Emotional Multimodal Actors Dataset
abstract
People convey their emotional state in their face and voice. We present an audio-visual data set uniquely suited for the study of multi-modal emotion expression and perception. The data set consists of facial and vocal emotional expressions in sentences spoken in a range of basic emotional states (happy, sad, anger, fear, disgust, and neutral). 7,442 clips of 91 actors with diverse ethnic backgrounds were rated by multiple raters in three modalities: audio, visual, and audio-visual. Categorical emotion labels and real-value intensity values for the perceived emotion were collected using crowd-sourcing from 2,443 raters. The human recognition of intended emotion for the audio-only, visual-only, and audio-visual data are 40.9%, 58.2% and 63.6% respectively. Recognition rates are highest for neutral, followed by happy, anger, disgust, fear, and sad. Average intensity levels of emotion are rated highest for visual-only perception. The accurate recognition of disgust and fear requires simultaneous audio-visual cues, while anger and happiness can be well recognized based on evidence from a single modality. The large dataset we introduce can be used to probe other questions concerning the audio-visual perception of emotion.
Houwei Cao, David G. Cooper, Michael K. Keutmann, Ruben C. Gur, Ani Nenkova, Ragini Verma
IEEE Trans. Affect. Comput.5
2013 Action Unit Models of Facial Expression of Emotion in the Presence of Speech
abstract
Automatic recognition of emotion using facial expressions in the presence of speech poses a unique challenge because talking reveals clues for the affective state of the speaker but distorts the canonical expression of emotion on the face. We introduce a corpus of acted emotion expression where speech is either present (talking) or absent (silent). The corpus is uniquely suited for analysis of the interplay between the two conditions. We use a multimodal decision level fusion classifier to combine models of emotion from talking and silent faces as well as from audio to recognize five basic emotions: anger, disgust, fear, happy and sad. Our results strongly indicate that emotion prediction in the presence of speech from action unit facial features is less accurate when the person is talking. Modeling talking and silent expressions separately and fusing the two models greatly improves accuracy of prediction in the talking setting. The advantages are most pronounced when silent and talking face models are fused with predictions from audio features. In this multi-modal prediction both the combination of modalities and the separate models of talking and silent facial expression of emotion contribute to the improvement.
Miraj Shah, David G. Cooper, Houwei Cao, Ruben C. Gur, Ani Nenkova, Ragini Verma
ACII5
2013 Automatic human utility evaluation of ASR systems: does WER really predict performance?
abstract
International audience
Benoît Favre, Kyla Cheung, Siavash Kazemian, Adam Lee, Yang Liu 0004, Cosmin Munteanu, Ani Nenkova, Dennis Ochei, Gerald Penn, Stephen Tratz, Clare R. Voss, Frauke Zeller
INTERSPEECH7
2013 Automatically Assessing Machine Summary Content Without a Gold Standard
abstract
The most widely adopted approaches for evaluation of summary content follow some protocol for comparing a summary with gold-standard human summaries, which are traditionally called model summaries. This evaluation paradigm falls short when human summaries are not available and becomes less accurate when only a single model is available. We propose three novel evaluation techniques. Two of them are model-free and do not rely on a gold standard for the assessment. The third technique improves standard automatic evaluations by expanding the set of available model summaries with chosen system summaries. We show that quantifying the similarity between the source text and its summary with appropriately chosen measures produces summary scores which replicate human assessments accurately. We also explore ways of increasing evaluation quality when only one human model summary is available as a gold standard. We introduce pseudomodels, which are system summaries deemed to contain good content according to automatic evaluation. Combining the pseudomodels with the single human model to form the gold-standard leads to higher correlations with human judgments compared to using only the one available model. Finally, we explore the feasibility of another measure—similarity between a system summary and the pool of all other system summaries for the same input. This method of comparison with the consensus of systems produces impressively accurate rankings of system summaries, achieving correlation with human rankings above 0.9.
Annie Louis, Ani Nenkova
Comput. Linguistics2
2013 What Makes Writing Great? First Experiments on Article Quality Prediction in the Science Journalism Domain
abstract
Great writing is rare and highly admired. Readers seek out articles that are beautifully written, informative and entertaining. Yet information-access technologies lack capabilities for predicting article quality at this level. In this paper we present first experiments on article quality prediction in the science journalism domain. We introduce a corpus of great pieces of science journalism, along with typical articles from the genre. We implement features to capture aspects of great writing, including surprising, visual and emotional content, as well as general features related to discourse organization and sentence structure. We show that the distinction between great and typical articles can be detected fairly accurately, and that the entire spectrum of our features contribute to the distinction.
Annie Louis, Ani Nenkova
Trans. Assoc. Comput. Linguistics2
2012 Lexical Differences in Autobiographical Narratives from Schizophrenic Patients and Healthy Controls
Kai Hong, Christian G. Kohler, Mary E. March, Amber A. Parker, Ani Nenkova
EMNLP-CoNLL5
2012 A Coherence Model Based on Syntactic Patterns
Annie Louis, Ani Nenkova
EMNLP-CoNLL2
2012 Combining video, audio and lexical indicators of affect in spontaneous conversation via particle filtering
abstract
We present experiments on fusing facial video, audio and lexical indicators for affect estimation during dyadic conversations. We use temporal statistics of texture descriptors extracted from facial video, a combination of various acoustic features, and lexical features to create regression based affect estimators for each modality. The single modality regressors are then combined using particle filtering, by treating these independent regression outputs as measurements of the affect states in a Bayesian filtering framework, where previous observations provide prediction about the current state by means of learned affect dynamics. Tested on the Audio-visual Emotion Recognition Challenge dataset, our single modality estimators achieve substantially higher scores than the official baseline method for every dimension of affect. Our filtering-based multi-modality fusion achieves correlation performance of 0.344 (baseline: 0.136) and 0.280 (baseline: 0.096) for the fully continuous and word level sub challenges, respectively.
Arman Savran, Houwei Cao, Miraj Shah, Ani Nenkova, Ragini Verma
ICMI4
2012 Combining Ranking and Classification to Improve Emotion Recognition in Spontaneous Speech
abstract
We introduce a novel emotion recognition approach which integrates ranking models. The approach is speaker independent, yet it is designed to exploit information from utterances from the same speaker in the test set before making predictions. It achieves much higher precision in identifying emotional utterances than a conventional SVM classifier. Furthermore we test several possibilities for combining conventional classification and predictions based on ranking. All combinations improve overall prediction accuracy. All experiments are performed on the FAU AIBO database which contains realistic spontaneous emotional speech. Our best combination system achieves 6.6 % absolute improvement over the Interspeech 2009 emotion challenge baseline system on the 5-class classification tasks. Index Terms: emotion classification, ranking models, spontaneous speech
Houwei Cao, Ragini Verma, Ani Nenkova
INTERSPEECH3
2012 A corpus of general and specific sentences from news
Annie Louis, Ani Nenkova
LREC2
2012 Acoustic-Prosodic Entrainment and Social Behavior
Rivka Levitan, Agustín Gravano, Laura Willson, Stefan Benus, Julia Hirschberg, Ani Nenkova
HLT-NAACL6
2012 Animating synthetic dyadic conversations with variations based on context and agent attributes
abstract
ABSTRACT Conversations between two people are ubiquitous in many inhabited contexts. The kinds of conversations that occur depend on several factors, including the time, the location of the participating agents, the spatial relationship between the agents, and the type of conversation in which they are engaged. The statistical distribution of dyadic conversations among a population of agents will therefore depend on these factors. In addition, the conversation types, flow, and duration will depend on agent attributes such as interpersonal relationships, emotional state, personal priorities, and socio‐cultural proxemics. We present a framework for distributing conversations among virtual embodied agents in a real‐time simulation. To avoid generating actual language dialogues, we express variations in the conversational flow by using behavior trees implementing a set of conversation archetypes. The flow of these behavior trees depends in part on the agents' attributes and progresses based on parametrically estimated transitional probabilities. With the participating agents' state, a ‘smart event’ model steers the interchange to different possible outcomes as it executes. Example behavior trees are developed for two conversation archetypes: buyer–seller negotiations and simple asking–answering; the model can be readily extended to others. Because the conversation archetype is known to participating agents, they can animate their gestures appropriate to their conversational state. The resulting animated conversations demonstrate reasonable variety and variability within the environmental context. Copyright © 2012 John Wiley & Sons, Ltd.
Libo Sun 0001, Alexander Shoulson, Nicole Nelson, Wenhu Qin, Ani Nenkova, Norman I. Badler
Comput. Animat. Virtual Worlds6
2011 Automatic identification of general and specific sentences by leveraging discourse annotations
Annie Louis, Ani Nenkova
IJCNLP2
2011 Acoustic and Prosodic Correlates of Social Behavior
abstract
We describe acoustic/prosodic and lexical correlates of social variables annotated on a large corpus of task-oriented spontaneous speech.We employ Amazon Mechanical Turk to label the corpus with a large number of social behaviors, examining results of three of these here.We find significant differences between male and female speakers for perceptions of attempts to be liked, likeability, speech planning, that also differ depending upon the gender of their conversational partners.
Agustín Gravano, Rivka Levitan, Laura Willson, Stefan Benus, Julia Hirschberg, Ani Nenkova
INTERSPEECH6
2011 Information Status Distinctions and Referring Expressions: An Empirical Study of References to People in News Summaries
abstract
Although there has been much theoretical work on using various information status distinctions to explain the form of references in written text, there have been few studies that attempt to automatically learn these distinctions for generating references in the context of computer-regenerated text. In this article, we present a model for generating references to people in news summaries that incorporates insights from both theory and a corpus analysis of human written summaries. In particular, our model captures how two properties of a person referred to in the summary—familiarity to the reader and global salience in the news story—affect the content and form of the initial reference to that person in a summary. We demonstrate that these two distinctions can be learned from a typical input for multi-document summarization and that they can be used to make regeneration decisions that improve the quality of extractive summaries.
Advaith Siddharthan, Ani Nenkova, Kathy McKeown
Comput. Linguistics2
2010 Automatic Evaluation of Linguistic Quality in Multi-Document Summarization
Emily Pitler, Annie Louis, Ani Nenkova
ACL3
2010 Creating Local Coherence: An Empirical Assessment
Annie Louis, Ani Nenkova
HLT-NAACL2
2010 Discourse indicators for content selection in summarization
Annie Louis, Aravind K. Joshi, Ani Nenkova
SIGDIAL Conference3
2010 Using entity features to classify implicit discourse relations
Annie Louis, Aravind K. Joshi, Rashmi Prasad, Ani Nenkova
SIGDIAL Conference4
2010 Class-level spectral features for emotion recognition
Dmitri Bitouk, Ragini Verma, Ani Nenkova
Speech Commun.3
2009 Automatic sense prediction for implicit discourse relations in text
Emily Pitler, Annie Louis, Ani Nenkova
ACL/IJCNLP3
2009 Predicting the Fluency of Text with Shallow Structural Features: Case Studies of Machine Translation and Human-Written Text
Jieun Chae, Ani Nenkova
EACL2
2009 Performance Confidence Estimation for Automatic Summarization
Annie Louis, Ani Nenkova
EACL2
2009 Automatically Evaluating Content Selection in Summarization without Human Models
Annie Louis, Ani Nenkova
EMNLP2
2009 Improving emotion recognition using class-level spectral features
abstract
Traditional approaches to automatic emotion recognition from speech typically make use of utterance level prosodic features. Still, a great deal of useful information about expressivity and emotion can be gained from segmental spectral features, which provide a more detailed description of the speech signal, or from measurements from specific regions of the utterance, such as the stressed vowels. Here we introduce a novel set of spectral features for emotion recognition: statistics of Mel-Frequency Spectral Coefficients computed over three phoneme type classes of interest: stressed vowels, unstressed vowels and consonants in the utterance. We investigate performance of our features in the task of speaker-independent emotion recognition using two publicly available datasets. Our experimental results clearly indicate that indeed both the richer set of spectral features and the differentiation between phoneme type classes are beneficial for the task. Classification accuracies are consistently higher for our features compared to prosodic features or utterance-level spectral features. Combination of our phoneme class features with prosodic features leads to even further improvement. Index Terms: emotion recognition
Dmitri Bitouk, Ani Nenkova, Ragini Verma
INTERSPEECH2
2008 Can You Summarize This? Identifying Correlates of Input Difficulty for Multi-Document Summarization
Ani Nenkova, Annie Louis
ACL1
2008 Revisiting Readability: A Unified Framework for Predicting Text Quality
Emily Pitler, Ani Nenkova
EMNLP2
2008 Entity-driven Rewrite for Multi-document Summarization
Ani Nenkova
IJCNLP1
2007 Measuring Importance and Query Relevance in Topic-focused Multi-document Summarization
Ani Nenkova, Daniel Jurafsky
ACL2
2007 Automatic detection of contrastive elements in spontaneous speech
abstract
In natural speech people use different levels of prominence to signal which parts of an utterance are especially important. Contrastive elements are often produced with stronger than usual prominence and their presence modifies the meaning of the utterance in subtle but important ways. We use a richly annotated corpus of conversational speech to study the acoustic characteristics of contrastive elements and the differences between them and words at other levels of prominence. We report our results for automatic detection of contrastive elements based on acoustic and textual features, finding that a baseline predicting nouns and adjectives as contrastive performs on par with the best combination of features. We achieve a much better performance in a modified task of detecting contrastive elements among words that are predicted to bear pitch accent.
Ani Nenkova, Daniel Jurafsky
ASRU1
2007 Modelling prominence and emphasis improves unit-selection synthesis
abstract
We describe the results of large scale perception experiments showing improvements in synthesising two distinct kinds of prominence: standard pitch-accent and strong emphatic accents. Previously prominence assignment has been mainly evaluated by computing accuracy on a prominence-labelled test set. By contrast we integrated an automatic pitch-accent classifier into the unit selection target cost and showed that listeners preferred these synthesised sentences. We also describe an improved recording script for collecting emphatic accents, and show that generating emphatic accents leads to further improvements in the fiction genre over incorporating pitch accent only. Finally, we show differences in the effects of prominence between child-directed speech and news and fiction genres. Index Terms: speech synthesis, prosody, prominence, pitch accent, unit selection
Volker Strom, Ani Nenkova, Robert A. J. Clark, Yolanda Vazquez-Alvarez, Jason M. Brenier, Simon King 0001, Daniel Jurafsky
INTERSPEECH2
2007 To Memorize or to Predict: Prominence labeling in Conversational Speech
Ani Nenkova, Jason M. Brenier, Anubha Kothari, Sasha Calhoun, Laura Whitton, David Beaver, Daniel Jurafsky
HLT-NAACL1
2007 Beyond SumBasic: Task-focused summarization with sentence simplification and lexical expansion
Lucy Vanderwende, Hisami Suzuki, Chris Brockett, Ani Nenkova
Inf. Process. Manag.4
2006 Summarization evaluation for text and speech: issues and approaches
abstract
This paper surveys current text and speech summarization evaluation approaches. It discusses advantages and disadvantages of these, with the goal of identifying summarization techniques most suitable to speech summarization. Precision/recall schemes, as well as summary accuracy measures which incorporate weightings based on multiple human decisions, are suggested as particularly suitable in evaluating speech summaries.
Ani Nenkova
INTERSPEECH1
2006 A compositional context sensitive multi-document summarizer: exploring the factors that influence summarization
abstract
The usual approach for automatic summarization is sentence extraction, where key sentences from the input documents are selected based on a suite of features. While word frequency often is used as a feature in summarization, its impact on system performance has not been isolated. In this paper, we study the contribution to summarization of three factors related to frequency: content word frequency, composition functions for estimating sentence importance from word frequency, and adjustment of frequency weights based on context. We carry out our analysis using datasets from the Document Understanding Conferences, studying not only the impact of these features on automatic summarizers, but also their role in human summarization. Our research shows that a frequency based summarizer can achieve performance comparable to that of state-of-the-art systems, but only with a good composition function; context sensitivity improves performance and significantly reduces repetition.
Ani Nenkova, Lucy Vanderwende, Kathy McKeown
SIGIR1
2006 The (Non)Utility of Linguistic Features for Predicting prominence in spontaneous speech
abstract
Conversational speech is characterized by prosodic variability which makes pitch accent prediction for this genre especially difficult. The linguistic literature points out that complex features such as information status, contrast and animacy help predict pitch accent placement. In this paper, we use a corpus annotated for such features to determine if they improve prominence prediction over traditional shallow features such as frequency and part-of-speech, or over new ones that we introduce. We demonstrate that while correlated with prominence, complex linguistic features do not improve prediction accuracy. Furthermore, the performance of our classifier is quite close to the ceiling defined by variability in human accent placement. An oracle experiment demonstrates, though, that at least some accuracy improvement is still possible.
Jason M. Brenier, Ani Nenkova, Anubha Kothari, Laura Whitton, David Beaver, Daniel Jurafsky
SLT2
2005 Automatic Text Summarization of Newswire: Lessons Learned from the Document Understanding Conference
Ani Nenkova
AAAI1
2005 Discourse Factors in Multi-Document Summarization
Ani Nenkova
AAAI1
2005 Do summaries help?
abstract
We describe a task-based evaluation to determine whether multi-document summaries measurably improve user performance whe using online news browsing systems for directed research. We evaluated the multi-document summaries generated by Newsblaster, a robust news browsing system that clusters online news articles and summarizes multiple articles on each event. Four groups of subjects were asked to perform the same time-restricted fact-gathering tasks, reading news under different conditions: no summaries at all, single sentence summaries drawn from one of the articles, Newsblaster multi-document summaries, and human summaries. Our results show that, in comparison to source documents only, the quality of reports assembled using Newsblaster summaries was significantly better and user satisfaction was higher with both Newsblaster and human summaries.
Kathy McKeown, Rebecca J. Passonneau, David K. Elson, Ani Nenkova, Julia Hirschberg
SIGIR4
2004 Syntactic Simplification for Improving Content Selection in Multi-Document Summarization
Advaith Siddharthan, Ani Nenkova, Kathy McKeown
COLING2
2004 Evaluating Content Selection in Summarization: The Pyramid Method
Ani Nenkova, Rebecca J. Passonneau
HLT-NAACL1
2003 Columbia's Newsblaster: New Features and Future Directions
Kathy McKeown, Regina Barzilay, David K. Elson, David Kirk Evans, Judith L. Klavans, Ani Nenkova, Barry Schiffman, Sergey Sigelman
HLT-NAACL7
2003 References to Named Entities: a Corpus Study
Ani Nenkova, Kathy McKeown
HLT-NAACL1