EDBT 2026 Demo / reviewers in the wild / expert
Eduard H. Hovy
dblp:47/2454
· DBLP profile ↗
224ranked-venue papers
22as first author
37since 2021 · last 2026
0000-0002-3270-7903ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 208 · 20 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 16 · 3 first-authorHuman-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Control Illusion: The Failure of Instruction Hierarchies in Large Language ModelsabstractLarge language models (LLMs) are increasingly deployed with hierarchical instruction schemes, where certain instructions (e.g., system-level directives) are expected to take precedence over others (e.g., user messages). Yet, we lack a systematic understanding of how effectively these hierarchical control mechanisms work. We introduce a systematic evaluation framework based on constraint prioritization to assess how well LLMs enforce instruction hierarchies. Our experiments across six state-of-the-art LLMs reveal that models struggle with consistent instruction prioritization, even for simple formatting conflicts. We find that the widely-adopted system/user prompt separation fails to establish a reliable instruction hierarchy, and models exhibit strong inherent biases toward certain constraint types regardless of their priority designation. Interestingly, we also find that societal hierarchy framings (e.g., authority, expertise, consensus) show stronger influence on model behavior than system/user roles, suggesting that pretraining-derived social structures function as latent behavioral priors with potentially greater impact than post-training guardrails. Yilin Geng 0001, Haonan Li 0002, Honglin Mu, Timothy Baldwin, Omri Abend, Eduard H. Hovy, Lea Frermann |
AAAI | 7 |
| 2026 | EVOTOOL: Self-Evolving Tool-Use Policy Optimization in LLM Agents via Blame-Aware Mutation and Diversity-Aware SelectionabstractShuo Yang, Caren Han, Xueqi Ma, Yan Li, Mohammad Reza Ghasemi Madani, Eduard Hovy. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Soyeon Caren Han, Xueqi Ma, Yan Li 0186, Mohammad Reza Ghasemi Madani, Eduard H. Hovy |
ACL (1) | 6 |
| 2026 | CommonMorph: Participatory Morphological Documentation Platform
Aso Mahmudi, Sina Ahmadi, Kemal Kurniawan, Rico Sennrich, Eduard H. Hovy, Ekaterina Vylomova |
LREC | 5 |
| 2026 | Retain or Reframe? A Computational Framework for the Analysis of Framing in News Articles and Reader CommentsabstractAbstract When a news article describes immigration as an “economic burden” or a “humanitarian crisis,” it selectively emphasizes certain aspects of the issue. Although this framing shapes how the public interprets such issues, audiences do not absorb frames passively but actively reorganize the presented information. While this relationship between source content and audience response is well-documented in the social sciences, NLP approaches often ignore it, analyzing frames in articles and responses in isolation. We present the first computational framework for large-scale analysis of framing across source content (news articles) and audience responses (reader comments). Methodologically, we refine frame labels and develop a framework that reconstructs primary frames in articles and comments from sentence-level predictions, and aligns articles with topically relevant comments. Applying our framework across eleven topics and two news outlets, we find that frame reuse in comments correlates highly across outlets, and that readers often selectively engage with frames of the articles. We release a frame classifier that performs well on both articles and comments, a dataset of article and comment sentences manually labeled for frames, and a large-scale dataset of articles and comments with predicted frame labels.1 Matteo Guida, Yulia Otmakhova 0001, Eduard H. Hovy, Lea Frermann |
Trans. Assoc. Comput. Linguistics | 3 |
| 2025 | RAEmoLLM: Retrieval Augmented LLMs for Cross-Domain Misinformation Detection Using In-Context Learning Based on Emotional InformationabstractMisinformation is prevalent in various fields such as education, politics, health, etc., causing significant harm to society. However, current methods for cross-domain misinformation detection rely on effort- and resource-intensive fine-tuning and complex model structures. With the outstanding performance of LLMs, many studies have employed them for misinformation detection. Unfortunately, they focus on in-domain tasks and do not incorporate significant sentiment and emotion features (which we jointly call affect). In this paper, we propose RAEmoLLM, the first retrieval augmented (RAG) LLMs framework to address cross-domain misinformation detection using in-context learning based on affective information. RAEmoLLM includes three modules. (1) In the index construction module, we apply an emotional LLM to obtain affective embeddings from all domains to construct a retrieval database. (2) The retrieval module uses the database to recommend top K examples (text-label pairs) from source domain data for target domain contents. (3) These examples are adopted as few-shot demonstrations for the inference module to process the target domain content. The RAEmoLLM can effectively enhance the general performance of LLMs in cross-domain misinformation detection tasks through affect-based retrieval, without fine-tuning. We evaluate our framework on three misinformation benchmarks. Results show that RAEmoLLM achieves significant improvements compared to the other few-shot methods on three datasets, with the highest increases of 15.64%, 31.18%, and 15.73% respectively. This project is available at https://github.com/lzw108/RAEmoLLM. Zhiwei Liu 0003, Kailai Yang, Qianqian Xie, Christine de Kock, Sophia Ananiadou, Eduard H. Hovy |
ACL (1) | 6 |
| 2025 | What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence FunctionsabstractLarge language models (LLMs) are trained on a vast amount of human-written data, but data providers often remain uncredited. In response to this issue, data valuation (or data attribution), which quantifies the contribution or value of each data to the model output, has been discussed as a potential solution. Nevertheless, applying existing data valuation methods to recent LLMs and their vast training datasets has been largely limited by prohibitive compute and memory costs. In this work, we focus on influence functions, a popular gradient-based data valuation method, and significantly improve its scalability with an efficient gradient projection strategy called LoGra that leverages the gradient structure in backpropagation. We then provide a theoretical motivation of gradient projection approaches to influence functions to promote trust in the data valuation process. Lastly, we lower the barrier to implementing data valuation systems by introducing LogIX, a software package that can transform existing training code into data valuation code with minimal effort. In our data valuation experiments, LoGra achieves competitive accuracy against more expensive baselines while showing up to 6,500x improvement in throughput and 5x reduction in GPU memory usage when applied to Llama3-8B-Instruct and the 1B-token dataset. Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff G. Schneider, Eduard H. Hovy, Roger B. Grosse, Eric P. Xing |
NeurIPS | 11 |
| 2024 | A Sentiment Consolidation Framework for Meta-Review GenerationabstractModern natural language generation systems with Large Language Models (LLMs) exhibit the capability to generate a plausible summary of multiple documents; however, it is uncertain if they truly possess the capability of information consolidation to generate summaries, especially on documents with opinionated information.We focus on meta-review generation, a form of sentiment summarisation for the scientific domain.To make scientific sentiment summarization more grounded, we hypothesize that human meta-reviewers follow a three-layer framework of sentiment consolidation to write meta-reviews.Based on the framework, we propose novel prompting methods for LLMs to generate meta-reviews and evaluation metrics to assess the quality of generated meta-reviews.Our framework is validated empirically as we find that prompting LLMs based on the framework -compared with prompting them with simple instructions -generates better metareviews.11 The code and annotated data are accessible at https: //github.com/oaimli/MetaReviewingLogic. Jey Han Lau, Eduard H. Hovy |
ACL (1) | 3 |
| 2024 | Transitive Consistency Constrained Learning for Entity-to-Entity Stance DetectionabstractEntity-to-entity stance detection identifies the stance between a pair of entities with a directed link that indicates the source, target and polarity.It is a streamlined task without the complex dependency structure for structural sentiment analysis, while it is more informative compared to most previous work assuming that the source is the author.Previous work performs entity-to-entity stance detection training on individual entity pairs.However, stances between inter-connected entity pairs may be correlated.In this paper, we propose transitive consistency constrained learning, which first finds connected entity pairs and their stances, and adds an additional objective to enforce the transitive consistency.We explore consistency training on both classification-based and generation-based models and conduct experiments to compare consistency training with previous work and large language models with in-context learning.Experimental results illustrate that the inter-correlation of stances in political news can be used to improve the entityto-entity stance detection model, while overly strict consistency enforcement may have a negative impact.In addition, we find that large language models struggle with predicting link direction and neutral labels in this task. 1 Haoyang Wen, Eduard H. Hovy, Alex Hauptmann 0001 |
ACL (1) | 2 |
| 2024 | DualGCN: Exploring Syntactic and Semantic Information for Aspect-Based Sentiment AnalysisabstractThe task of aspect-based sentiment analysis aims to identify sentiment polarities of given aspects in a sentence. Recent advances have demonstrated the advantage of incorporating the syntactic dependency structure with graph convolutional networks (GCNs). However, their performance of these GCN-based methods largely depends on the dependency parsers, which would produce diverse parsing results for a sentence. In this article, we propose a dual GCN (DualGCN) that jointly considers the syntax structures and semantic correlations. Our DualGCN model mainly comprises four modules: 1) SynGCN: instead of explicitly encoding syntactic structure, the SynGCN module uses the dependency probability matrix as a graph structure to implicitly integrate the syntactic information; 2) SemGCN: we design the SemGCN module with multihead attention to enhance the performance of the syntactic structure with the semantic information; 3) Regularizers: we propose orthogonal and differential regularizers to precisely capture semantic correlations between words by constraining attention scores in the SemGCN module; and 4) Mutual BiAffine: we use the BiAffine module to bridge relevant information between the SynGCN and SemGCN modules. Extensive experiments are conducted compared with up-to-date pretrained language encoders on two groups of datasets, one including Restaurant14, Laptop14, and Twitter and the other including Restaurant15 and Restaurant16. The experimental results demonstrate that the parsing results of various dependency parsers affect their performance of the GCN-based models. Our DualGCN model achieves superior performance compared with the state-of-the-art approaches. The source code and preprocessed datasets are provided and publicly available on GitHub (see https://github.com/CCChenhao997/DualGCN-ABSA). Ruifan Li, Hao Chen 0041, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006, Eduard H. Hovy |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | A Textual Dataset for Situated Proactive Response SelectionabstractRecent data-driven conversational models are able to return fluent, consistent, and informative responses to many kinds of requests and utterances in task-oriented scenarios.However, these responses are typically limited to just the immediate local topic instead of being widerranging and proactively taking the conversation further, for example making suggestions to help customers achieve their goals.This inadequacy reflects a lack of understanding of the interlocutor's situation and implicit goal.To address the problem, we introduce a task of proactive response selection based on situational information.We present a manuallycurated dataset of 1.7k English conversation examples that include situational background information plus for each conversation a set of responses, only some of which are acceptable in the situation.A responsive and informed conversation system should select the appropriate responses and avoid inappropriate ones; doing so demonstrates the ability to adequately understand the initiating request and situation.Our benchmark experiments show that this is not an easy task even for strong neural models, offering opportunities for future research. Naoki Otani, Jun Araki, HyeongSik Kim 0001, Eduard H. Hovy |
ACL (1) | 4 |
| 2023 | What's the Meaning of Superhuman Performance in Today's NLU?abstractSimone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajič, Daniel Hershcovich, Eduard Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, Roberto Navigli. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajic 0001, Daniel Hershcovich, Eduard H. Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, Roberto Navigli |
ACL (1) | 6 |
| 2023 | CHARD: Clinical Health-Aware Reasoning Across Dimensions for Text Generation ModelsabstractWe motivate and introduce CHARD: Clinical Health-Aware Reasoning across Dimensions, to investigate the capability of text generation models to act as implicit clinical knowledge bases and generate free-flow textual explanations about various health-related conditions across several dimensions.We collect and present an associated dataset, CHARDat, consisting of explanations about 52 health conditions across three clinical dimensions.We conduct extensive experiments using BART and T5 along with data augmentation, and perform automatic, human, and qualitative analyses.We show that while our models can perform decently, CHARD is very challenging with strong potential for further exploration. Steven Y. Feng, Vivek Khetan, Bogdan Sacaleanu, Anatole Gershman, Eduard H. Hovy |
EACL | 5 |
| 2023 | PANCETTA: Phoneme Aware Neural Completion to Elicit Tongue Twisters AutomaticallyabstractTongue twisters are meaningful sentences that are difficult to pronounce.The process of automatically generating tongue twisters is challenging since the generated utterance must satisfy two conditions at once: phonetic difficulty and semantic meaning.Furthermore, phonetic difficulty is itself hard to characterize and is expressed in tongue twisters through a heterogeneous mix of phenomena such as alliteration and homophony.In this paper, we propose PANCETTA: Phoneme Aware Neural Completion to Elicit Tongue Twisters Automatically.We leverage phoneme representations to capture the notion of phonetic difficulty, and we train language models to generate original tongue twisters on two proposed task settings.To do this, we curate a dataset called TT-Corp, consisting of existing English tongue twisters.Through automatic and human evaluation, as well as qualitative analysis, we show that PANCETTA generates novel, phonetically difficult, fluent, and semantically meaningful tongue twisters. Sedrick Keh, Steven Y. Feng, Varun Gangal, Malihe Alikhani, Eduard H. Hovy |
EACL | 5 |
| 2022 | Retrieve, Caption, Generate: Visual Grounding for Enhancing Commonsense in Text Generation ModelsabstractWe investigate the use of multimodal information contained in images as an effective method for enhancing the commonsense of Transformer models for text generation. We perform experiments using BART and T5 on concept-to-text generation, specifically the task of generative commonsense reasoning, or CommonGen. We call our approach VisCTG: Visually Grounded Concept-to-Text Generation. VisCTG involves captioning images representing appropriate everyday scenarios, and using these captions to enrich and steer the generation process. Comprehensive evaluation and analysis demonstrate that VisCTG noticeably improves model performance while successfully addressing several issues of the baseline generations, including poor commonsense, fluency, and specificity. Steven Y. Feng, Zhuofu Tao, Malihe Alikhani, Teruko Mitamura, Eduard H. Hovy, Varun Gangal |
AAAI | 6 |
| 2022 | NAREOR: The Narrative Reordering ProblemabstractMany implicit inferences exist in text depending on how it is structured that can critically impact the text's interpretation and meaning. One such structural aspect present in text with chronology is the order of its presentation. For narratives or stories, this is known as the narrative order. Reordering a narrative can impact the temporal, causal, event-based, and other inferences readers draw from it, which in turn can have strong effects both on its interpretation and interestingness. In this paper, we propose and investigate the task of Narrative Reordering (NAREOR) which involves rewriting a given story in a different narrative order while preserving its plot. We present a dataset, NAREORC, with human rewritings of stories within ROCStories in non-linear orders, and conduct a detailed analysis of it. Further, we propose novel task-specific training methods with suitable evaluation metrics. We perform experiments on NAREORC using state-of-the-art models such as BART and T5 and conduct extensive automatic and human evaluations. We demonstrate that although our models can perform decently, NAREOR is a challenging task with potential for further exploration. We also investigate two applications of NAREOR: generation of more interesting variations of stories and serving as adversarial sets for temporal/event-related tasks, besides discussing other prospective ones, such as for pedagogical setups related to language skills like essay writing and applications to medicine involving clinical narratives. Varun Gangal, Steven Y. Feng, Malihe Alikhani, Teruko Mitamura, Eduard H. Hovy |
AAAI | 5 |
| 2022 | PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personification Data for Learning Enhanced GenerationabstractA personification is a figure of speech that endows inanimate entities with properties and actions typically seen as requiring animacy. In this paper, we explore the task of personification generation. To this end, we propose PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personification data for Learning Enhanced generation. We curate a corpus of personifications called PersonifCorp, together with automatically generated de-personified literalizations of these personifications. We demonstrate the usefulness of this parallel corpus by training a seq2seq model to personify a given literal input. Both automatic and human evaluations show that fine-tuning with PersonifCorp leads to significant gains in personification-related qualities such as animacy and interestingness. A detailed qualitative analysis also highlights key strengths and imperfections of PINEAPPLE over baselines, demonstrating a strong ability to generate diverse and creative personifications that enhance the overall appeal of a sentence. Sedrick Keh, Varun Gangal, Steven Y. Feng, Harsh Jhamtani, Malihe Alikhani, Eduard H. Hovy |
COLING | 7 |
| 2022 | NewsClaims: A New Benchmark for Claim Detection from News with Attribute KnowledgeabstractRevanth Gangi Reddy, Sai Chetan Chinthakindi, Zhenhailong Wang, Yi Fung, Kathryn Conger, Ahmed ELsayed, Martha Palmer, Preslav Nakov, Eduard Hovy, Kevin Small, Heng Ji. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Revanth Gangi Reddy, Sai Chetan Chinthakindi, Zhenhailong Wang, Yi R. Fung 0001, Kathryn Conger, Ahmed Elsayed, Martha Palmer, Preslav Nakov, Eduard H. Hovy, Kevin Small, Heng Ji 0001 |
EMNLP | 9 |
| 2022 | EvEntS ReaLM: Event Reasoning of Entity States via Language ModelsabstractThis paper investigates models of event implications.Specifically, how well models predict entity state-changes, by targeting their understanding of physical attributes.Nominally, Large Language models (LLM) have been exposed to procedural knowledge about how objects interact, yet our benchmarking shows they fail to reason about the world.Conversely, we also demonstrate that existing approaches often misrepresent the surprising abilities of LLMs via improper task encodings and that proper model prompting can dramatically improve performance of reported baseline results across multiple tasks.In particular, our results indicate that our prompting technique is especially useful for unseen attributes (out-of-domain) or when only limited data is available.1 Evangelia Spiliopoulou, Artidoro Pagnoni, Yonatan Bisk, Eduard H. Hovy |
EMNLP | 4 |
| 2022 | Transfer Learning from Semantic Role Labeling to Event Argument Extraction with Template-based Slot QueryingabstractIn this work, we investigate transfer learning from semantic role labeling (SRL) to event argument extraction (EAE), considering their similar argument structures.We view the extraction task as a role querying problem, unifying various methods into a single framework.There are key discrepancies on role labels and distant arguments between semantic role and event argument annotations.To mitigate these discrepancies, we specify natural language-like queries to tackle the label mismatch problem and devise argument augmentation to recover distant arguments.We show that SRL annotations can serve as a valuable resource for EAE, and a template-based slot querying strategy is especially effective for facilitating the transfer.In extensive evaluations on two English EAE benchmarks, our proposed model obtains impressive zero-shot results by leveraging SRL annotations, reaching nearly 80% of the fullysupervised scores.It further provides benefits in low-resource cases, where few EAE annotations are available.Moreover, we show that our approach generalizes to cross-domain and multilingual scenarios. Zhisong Zhang, Emma Strubell, Eduard H. Hovy |
EMNLP | 3 |
| 2022 | A Survey of Active Learning for Natural Language ProcessingabstractIn this work, we provide a literature review of active learning (AL) for its applications in natural language processing (NLP).In addition to a fine-grained categorization of query strategies, we also investigate several other important aspects of applying AL to NLP problems.These include AL for structured prediction tasks, annotation cost, model learning (especially with deep neural models), and starting and stopping AL.Finally, we conclude with a discussion of related topics and future directions. Zhisong Zhang, Emma Strubell, Eduard H. Hovy |
EMNLP | 3 |
| 2022 | One Document, Many Revisions: A Dataset for Classification and Description of Edit IntentsabstractDocument authoring involves a lengthy revision process, marked by individual edits that are frequently linked to comments. Modeling the relationship between edits and comments leads to a better understanding of document evolution, potentially benefiting applications such as content summarization, and task triaging. Prior work on understanding revisions has primarily focused on classifying edit intents, but falling short of a deeper understanding of the nature of these edits. In this paper, we present explore the challenge of describing an edit at two levels: identifying the edit intent, and describing the edit using free-form text. We begin by defining a taxonomy of general edit intents and introduce a new dataset of full revision histories of Wikipedia pages, annotated with each revision’s edit intent. Using this dataset, we train a classifier that achieves a 90% accuracy in identifying edit intent. We use this classifier to train a distantly-supervised model that generates a high-level description of a revision in free-form text. Our experimental results show that incorporating edit intent information aids in generating better edit descriptions. We establish a set of baselines for the edit description task, achieving a best score of 28 ROUGE, thus demonstrating the effectiveness of our layered approach to edit understanding. Dheeraj Rajagopal, Xuchao Zhang, Michael Gamon, Sujay Kumar Jauhar, Diyi Yang, Eduard H. Hovy |
LREC | 6 |
| 2021 | More Identifiable yet Equally Performant Transformers for Text ClassificationabstractRishabh Bhardwaj, Navonil Majumder, Soujanya Poria, Eduard Hovy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Rishabh Bhardwaj, Navonil Majumder, Soujanya Poria, Eduard H. Hovy |
ACL/IJCNLP (1) | 4 |
| 2021 | Style is NOT a single variable: Case Studies for Cross-Stylistic Language UnderstandingabstractDongyeop Kang, Eduard Hovy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Dongyeop Kang, Eduard H. Hovy |
ACL/IJCNLP (1) | 2 |
| 2021 | Dual Graph Convolutional Networks for Aspect-based Sentiment AnalysisabstractRuifan Li, Hao Chen, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang, Eduard Hovy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ruifan Li, Hao Chen 0041, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006, Eduard H. Hovy |
ACL/IJCNLP (1) | 6 |
| 2021 | Probing the Probing Paradigm: Does Probing Accuracy Entail Task Relevance?abstractAlthough neural models have achieved impressive results on several NLP benchmarks, little is understood about the mechanisms they use to perform language tasks.Thus, much recent attention has been devoted to analyzing the sentence representations learned by neural encoders, through the lens of 'probing' tasks.However, to what extent was the information encoded in sentence representations, as discovered through a probe, actually used by the model to perform its task?In this work, we examine this probing paradigm through a case study in Natural Language Inference, showing that models can learn to encode linguistic properties even if they are not needed for the task on which the model was trained.We further identify that pretrained word embeddings play a considerable role in encoding these properties rather than the training task itself, highlighting the importance of careful controls when designing probing experiments.Finally, through a set of controlled synthetic tasks, we demonstrate models can encode these properties considerably above chance-level even when distributed in the data as random noise, calling into question the interpretation of absolute claims on probing tasks. 1 Abhilasha Ravichander, Yonatan Belinkov, Eduard H. Hovy |
EACL | 3 |
| 2021 | NoiseQA: Challenge Set Evaluation for User-Centric Question AnsweringabstractAbhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard Hovy, Alan W Black. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Abhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard H. Hovy, Alan W. Black |
EACL | 5 |
| 2021 | Investigating Robustness of Dialog Models to Popular Figurative Language ConstructsabstractHumans often employ figurative language use in communication, including during interactions with dialog systems.Thus, it is important for real-world dialog systems to be able to handle popular figurative language constructs like metaphor and simile.In this work, we analyze the performance of existing dialog models in situations where the input dialog context exhibits use of figurative language.We observe large gaps in handling of figurative language when evaluating the models on two open domain dialog datasets.When faced with dialog contexts consisting of figurative language, some models show very large drops in performance compared to contexts without figurative language.We encourage future research in dialog modeling to separately analyze and report results on figurative language in order to better test model capabilities relevant to real-world use.Finally, we propose lightweight solutions to help existing models become more robust to figurative language by simply using an external resource to translate figurative language to literal (non-figurative) forms while preserving the meaning to the best extent possible. Harsh Jhamtani, Varun Gangal, Eduard H. Hovy, Taylor Berg-Kirkpatrick |
EMNLP (1) | 3 |
| 2021 | Think about it! Improving defeasible reasoning by first modeling the question scenarioabstractDefeasible reasoning is the mode of reasoning where conclusions can be overturned by taking into account new evidence.Existing cognitive science literature on defeasible reasoning suggests that a person forms a mental model of the problem scenario before answering questions.Our research goal asks whether neural models can similarly benefit from envisioning the question scenario before answering a defeasible query.Our approach is, given a question, to have a model first create a graph of relevant influences, and then leverage that graph as an additional input when answering the question.Our system, CURIOUS, achieves a new stateof-the-art on three different defeasible reasoning datasets.This result is significant as it illustrates that performance can be improved by guiding a system to "think about" a question and explicitly model the scenario, rather than answering reflexively. 1 Aman Madaan, Niket Tandon, Dheeraj Rajagopal, Peter Clark, Yiming Yang 0002, Eduard H. Hovy |
EMNLP (1) | 6 |
| 2021 | SELFEXPLAIN: A Self-Explaining Architecture for Neural Text ClassifiersabstractWe introduce SELFEXPLAIN, a novel selfexplaining model that explains a text classifier's predictions using phrase-based concepts.SELFEXPLAIN augments existing neural classifiers by adding (1) a globally interpretable layer that identifies the most influential concepts in the training set for a given sample and (2) a locally interpretable layer that quantifies the contribution of each local input concept by computing a relevance score relative to the predicted label.Experiments across five text-classification datasets show that SELFEX-PLAIN facilitates interpretability without sacrificing performance.Most importantly, explanations from SELFEXPLAIN show sufficiency for model predictions and are perceived as adequate, trustworthy and understandable by human judges compared to existing widely-used baselines.1 Dheeraj Rajagopal, Vidhisha Balachandran, Eduard H. Hovy, Yulia Tsvetkov |
EMNLP (1) | 3 |
| 2021 | On the Benefit of Syntactic Supervision for Cross-lingual Transfer in Semantic Role LabelingabstractAlthough recent developments in neural architectures and pre-trained representations have greatly increased state-of-the-art model performance on fully-supervised semantic role labeling (SRL), the task remains challenging for languages where supervised SRL training data are not abundant.Cross-lingual learning can improve performance in this setting by transferring knowledge from high-resource languages to low-resource ones.Moreover, we hypothesize that annotations of syntactic dependencies can be leveraged to further facilitate cross-lingual transfer.In this work, we perform an empirical exploration of the helpfulness of syntactic supervision for crosslingual SRL within a simple multitask learning scheme.With comprehensive evaluations across ten languages (in addition to English) and three SRL benchmark datasets, including both dependency-and span-based SRL, we show the effectiveness of syntactic supervision in low-resource scenarios.Experiments Target Languages SRL Style Same Frames?Compatible Roles?Main SRL Setting EWT/UPB † ( §3.2) de,fr,it,es,pt,fi Dependency-based Yes Yes Zero-shot EWT/FiPB ( §3.3) fi Dependency-based No Yes Semi-supervised CoNLL-2009 ( §3.4) cs,zh,es,ca Dependency-based No No Semi-supervised OntoNotes ( §3.5) zh,ar Span-based No Yes Semi-supervised Zhisong Zhang, Emma Strubell, Eduard H. Hovy |
EMNLP (1) | 3 |
| 2021 | Explaining the Efficacy of Counterfactually Augmented Data
Divyansh Kaushik, Amrith Setlur, Eduard H. Hovy, Zachary C. Lipton |
ICLR | 3 |
| 2021 | Decoupling Global and Local Representations via Invertible Generative Flows
Xuezhe Ma, Xiang Kong, Shanghang Zhang, Eduard H. Hovy |
ICLR | 4 |
| 2021 | SAPPHIRE: Approaches for Enhanced Concept-to-Text GenerationabstractWe motivate and propose a suite of simple but effective improvements for concept-to-text generation called SAPPHIRE: Set Augmentation and Post-hoc PHrase Infilling and REcombination.We demonstrate their effectiveness on generative commonsense reasoning, a.k.a. the CommonGen task, through experiments using both BART and T5 models.Through extensive automatic and human evaluation, we show that SAPPHIRE noticeably improves model performance.An in-depth qualitative analysis illustrates that SAPPHIRE effectively addresses many issues of the baseline model generations, including lack of commonsense, insufficient specificity, and poor fluency. Steven Y. Feng, Jessica Huynh, Chaitanya Narisetty, Eduard H. Hovy, Varun Gangal |
INLG | 4 |
| 2021 | StylePTB: A Compositional Benchmark for Fine-grained Controllable Text Style TransferabstractYiwei Lyu, Paul Pu Liang, Hai Pham, Eduard Hovy, Barnabás Póczos, Ruslan Salakhutdinov, Louis-Philippe Morency. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yiwei Lyu 0001, Paul Pu Liang, Hai Pham, Eduard H. Hovy, Barnabás Póczos, Ruslan Salakhutdinov, Louis-Philippe Morency |
NAACL-HLT | 4 |
| 2021 | Measuring and Improving Consistency in Pretrained Language ModelsabstractAbstract Consistency of a model—that is, the invariance of its behavior under meaning-preserving alternations in its input—is a highly desirable property in natural language processing. In this paper we study the question: Are Pretrained Language Models (PLMs) consistent with respect to factual knowledge? To this end, we create ParaRel🤘, a high-quality resource of cloze-style query English paraphrases. It contains a total of 328 paraphrases for 38 relations. Using ParaRel🤘, we show that the consistency of all PLMs we experiment with is poor— though with high variance between relations. Our analysis of the representational spaces of PLMs suggests that they have a poor structure and are currently not suitable for representing knowledge robustly. Finally, we propose a method for improving model consistency and experimentally demonstrate its effectiveness.1 Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, Yoav Goldberg |
Trans. Assoc. Comput. Linguistics | 5 |
| 2021 | Erratum: Measuring and Improving Consistency in Pretrained Language ModelsabstractAbstract During production of this paper, an error was introduced to the formula on the bottom of the right column of page 1020. In the last two terms of the formula, the n and m subscripts were swapped. The correct formula is:Lc=∑n=1k∑m=n+1kDKL(Qnri∥Qmri)+DKL(Qmri∥Qnri)The paper has been updated. Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard H. Hovy, Hinrich Schütze, Yoav Goldberg |
Trans. Assoc. Comput. Linguistics | 5 |
| 2021 | Classifying Argumentative Relations Using Logical Mechanisms and Argumentation SchemesabstractWhile argument mining has achieved significant success in classifying argumentative relations between statements (support, attack, and neutral), we have a limited computational understanding of logical mechanisms that constitute those relations. Most recent studies rely on black-box models, which are not as linguistically insightful as desired. On the other hand, earlier studies use rather simple lexical features, missing logical relations between statements. To overcome these limitations, our work classifies argumentative relations based on four logical and theory-informed mechanisms between two statements, namely, (i) factual consistency, (ii) sentiment coherence, (iii) causal relation, and (iv) normative relation. We demonstrate that our operationalization of these logical mechanisms classifies argumentative relations without directly training on data labeled with the relations, significantly better than several unsupervised baselines. We further demonstrate that these mechanisms also improve supervised classifiers through representation learning. Yohan Jo, Seo-Jin Bang, Chris Reed 0001, Eduard H. Hovy |
Trans. Assoc. Comput. Linguistics | 4 |
| 2020 | SCDE: Sentence Cloze Dataset with High Quality Distractors From ExaminationsabstractWe introduce SCDE, a dataset to evaluate the performance of computational models through sentence prediction.SCDE is a humancreated sentence cloze dataset, collected from public school English examinations.Our task requires a model to fill up multiple blanks in a passage from a shared candidate set with distractors designed by English teachers.Experimental results demonstrate that this task requires the use of non-local, discourse-level context beyond the immediate sentence neighborhood.The blanks require joint solving and significantly impair each other's context.Furthermore, through ablations, we show that the distractors are of high quality and make the task more challenging.Our experiments show that there is a significant performance gap between advanced models (72%) and humans (87%), encouraging future models to bridge this gap. 1 2 Passage: A student's life is never easy.And it is even more difficult if you will have to complete your study in a foreign land. 1The following are some basic things you need to do before even seizing that passport and boarding on the plane.Knowing the country.You shouldn't bother researching the country's hottest tourist spots or historical places.You won't go there as a tourist, but as a student.2 In addition, read about their laws.You surely don't want to face legal problems, especially if you're away from home.3 Don't expect that you can graduate abroad without knowing even the basics of the language.Before leaving your home country, take online lessons to at least master some of their words and sentences.This will be useful in living and studying there.Doing this will also prepare you in communicating with those who can't speak English.Preparing for other needs.Check the conversion of your money to their local currency.4. The Internet of your intended school will be very helpful in findings an apartment and helping you understand local currency.Remember, you're not only carrying your own reputation but your country's reputation as well.If you act foolishly, people there might think that all of your countrymen are foolish as well. 5Candidates: A. Studying their language.B. That would surely be a very bad start for your study abroad program.C. Going with their trends will keep it from being too obvious that you're a foreigner.D. Set up your bank account so you can use it there, get an insurance, and find an apartment.E. It'll be helpful to read the most important points in their history and to read up on their culture.F. A lot of preparations are needed so you can be sure to go back home with a diploma and a bright future waiting for you.G. Packing your clothes.Answers with Reasoning Type: 1→F (Summary) , 2→E (Inference) , 3→A (Paraphrase) , 4→D (WordMatch), 5→B (Inference) (C and G are distractors) Discussion: Blank 3 is the easiest to solve, since "Studying their language" is a near-paraphrase of "Knowing even the basics of the language".Blank 2 needs to be reasoned out by Inference -specifically E can be inferred from the previous sentence.Note however that C is also a possible inference from the previous sentence -it is only after reading the entire context, which seems to be about learning various aspects of a country, that E seems to fit better.Blank 1 needs Summary → it requires understanding several later sentences and abstracting out that they all refer to lots of preparations.Finally, Blank 5 can be mapped to B by inferring that people thinking all your countrymen are foolish is bad, while Blank 4 is a easy WordMatch on apartment to D. The other distractor G, although topically related to preparation for going abroad, does not directly fit into any of the blank contexts Xiang Kong, Varun Gangal, Eduard H. Hovy |
ACL | 3 |
| 2020 | A Two-Step Approach for Implicit Event Argument DetectionabstractIn this work, we explore the implicit event argument detection task, which studies event arguments beyond sentence boundaries.The addition of cross-sentence argument candidates imposes great challenges for modeling.To reduce the number of candidates, we adopt a two-step approach, decomposing the problem into two sub-problems: argument head-word detection and head-to-span expansion.Evaluated on the recent RAMS dataset (Ebner et al., 2020), our model achieves overall better performance than a strong sequence labeling baseline.We further provide detailed error analysis, presenting where the model mainly makes errors and indicating directions for future improvements.It remains a challenge to detect implicit arguments, calling for more future work of document-level modeling for this task. Zhisong Zhang, Xiang Kong, Zhengzhong Liu 0001, Xuezhe Ma, Eduard H. Hovy |
ACL | 5 |
| 2020 | Measuring Forecasting Skill from TextabstractPeople vary in their ability to make accurate predictions about the future.Prior studies have shown that some individuals can predict the outcome of future events with consistently better accuracy.This leads to a natural question: what makes some forecasters better than others?In this paper we explore connections between the language people use to describe their predictions and their forecasting skill.Datasets from two different forecasting domains are explored: (1) geopolitical forecasts from Good Judgment Open, an online prediction forum and (2) a corpus of company earnings forecasts made by financial analysts.We present a number of linguistic metrics which are computed over text associated with people's predictions about the future including: uncertainty, readability, and emotion.By studying linguistic factors associated with predictions, we are able to shed some light on the approach taken by skilled forecasters.Furthermore, we demonstrate that it is possible to accurately predict forecasting skill using a model that is based solely on language.This could potentially be useful for identifying accurate predictions or potentially skilled forecasters earlier. 1 Shi Zong, Alan Ritter, Eduard H. Hovy |
ACL | 3 |
| 2020 | Definition Frames: Using Definitions for Hybrid Concept RepresentationsabstractAdvances in word representations have shown tremendous improvements in downstream NLP tasks, but lack semantic interpretability.In this paper 1 , we introduce Definition Frames (DF), a matrix distributed representation extracted from definitions, where each dimension is semantically interpretable.DF dimensions correspond to the Qualia structure relations (Boguraev and Pustejovsky, 1990): a set of relations that uniquely define a term.Our results show that DFs have competitive performance with other distributional semantic approaches on word similarity tasks. Evangelia Spiliopoulou, Artidoro Pagnoni, Eduard H. Hovy |
COLING | 3 |
| 2020 | Self-Training With Noisy Student Improves ImageNet ClassificationabstractWe present a simple self-training method that achieves 88.4% top-1 accuracy on ImageNet, which is 2.0% better than the state-of-the-art model that requires 3.5B weakly labeled Instagram images. On robustness test sets, it improves ImageNet-A top-1 accuracy from 61.0% to 83.7%, reduces ImageNet-C mean corruption error from 45.7 to 28.3, and reduces ImageNet-P mean flip rate from 27.8 to 12.2. To achieve this result, we first train an EfficientNet model on labeled ImageNet images and use it as a teacher to generate pseudo labels on 300M unlabeled images. We then train a larger EfficientNet as a student model on the combination of labeled and pseudo labeled images. We iterate this process by putting back the student as the teacher. During the generation of the pseudo labels, the teacher is not noised so that the pseudo labels are as accurate as possible. However, during the learning of the student, we inject noise such as dropout, stochastic depth and data augmentation via RandAugment to the student so that the student generalizes better than the teacher. Qizhe Xie, Minh-Thang Luong, Eduard H. Hovy, Quoc V. Le |
CVPR | 3 |
| 2020 | Detecting Attackable Sentences in ArgumentsabstractFinding attackable sentences in an argument is the first step toward successful refutation in argumentation.We present a first large-scale analysis of sentence attackability in online arguments.We analyze driving reasons for attacks in argumentation and identify relevant characteristics of sentences.We demonstrate that a sentence's attackability is associated with many of these characteristics regarding the sentence's content, proposition types, and tone, and that an external knowledge source can provide useful information about attackability.Building on these findings, we demonstrate that machine learning models can automatically detect attackable sentences in arguments, significantly better than several baselines and comparably well to laypeople. 1 Yohan Jo, Seo-Jin Bang, Emaad Manzoor, Eduard H. Hovy, Chris Reed 0001 |
EMNLP (1) | 4 |
| 2020 | Extracting Implicitly Asserted Propositions in ArgumentationabstractArgumentation accommodates various rhetorical devices, such as questions, reported speech, and imperatives.These rhetorical tools usually assert argumentatively relevant propositions rather implicitly, so understanding their true meaning is key to understanding certain arguments properly.However, most argument mining systems and computational linguistics research have paid little attention to implicitly asserted propositions in argumentation.In this paper, we examine a wide range of computational methods for extracting propositions that are implicitly asserted in questions, reported speech, and imperatives in argumentation.By evaluating the models on a corpus of 2016 U.S. presidential debates and online commentary, we demonstrate the effectiveness and limitations of the computational models.Our study may inform future research on argument mining and the semantics of these rhetorical devices in argumentation. 1 Yohan Jo, Jacky Visser, Chris Reed 0001, Eduard H. Hovy |
EMNLP (1) | 4 |
| 2020 | Plan ahead: Self-Supervised Text Planning for Paragraph Completion TaskabstractDespite the recent success of contextualized language models on various NLP tasks, language model itself cannot capture textual coherence of a long, multi-sentence document (e.g., a paragraph).Humans often make structural decisions on what and how to say about before making utterances.Guiding surface realization with such high-level decisions and structuring text in a coherent way is essentially called a planning process.Where can the model learn such high-level coherence?A paragraph itself contains various forms of inductive coherence signals called self-supervision in this work, such as sentence orders, topical keywords, rhetorical structures, and so on.Motivated by that, this work proposes a new paragraph completion task PAR-COM; predicting masked sentences in a paragraph.However, the task suffers from predicting and selecting appropriate topical content with respect to the given context.To address that, we propose a self-supervised text planner SSPlanner that predicts what to say first (content prediction), then guides the pretrained language model (surface realization) using the predicted content.SSPlanner outperforms the baseline generation models on the paragraph completion task in both automatic and human evaluation.We also find that a combination of noun and verb types of keywords is the most effective for content selection.As more number of content keywords are provided, overall generation quality also increases. Dongyeop Kang, Eduard H. Hovy |
EMNLP (1) | 2 |
| 2020 | Incorporating a Local Translation Mechanism into Non-autoregressive TranslationabstractIn this work, we introduce a novel local autoregressive translation (LAT) mechanism into non-autoregressive translation (NAT) models so as to capture local dependencies among target outputs.Specifically, for each target decoding position, instead of only one token, we predict a short sequence of tokens in an autoregressive way.We further design an efficient merging algorithm to align and merge the output pieces into one final output sequence.We integrate LAT into the conditional masked language model (CMLM; Ghazvininejad et al., 2019) and similarly adopt iterative decoding.Empirical results on five translation tasks show that compared with CMLM, our method achieves comparable or better performance with fewer decoding iterations, bringing a 2.5x speedup.Further analysis indicates that our method reduces repeated translations and performs better at longer sentences.The code for our model is available at https://github. com/shawnkx/NAT-with-Local-AT. Xiang Kong, Zhisong Zhang, Eduard H. Hovy |
EMNLP (1) | 3 |
| 2020 | A Dataset for Tracking Entities in Open Domain Procedural TextabstractNiket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson, Eduard Hovy. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Niket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson 0001, Eduard H. Hovy |
EMNLP (1) | 8 |
| 2020 | Learning The Difference That Makes A Difference With Counterfactually-Augmented Data
Divyansh Kaushik, Eduard H. Hovy, Zachary C. Lipton |
ICLR | 2 |
| 2020 | Machine-Aided Annotation for Fine-Grained Proposition Types in ArgumentationabstractWe introduce a corpus of the 2016 U.S. presidential debates and commentary, containing 4,648 argumentative propositions annotated with fine-grained proposition types. Modern machine learning pipelines for analyzing argument have difficulty distinguishing between types of propositions based on their factuality, rhetorical positioning, and speaker commitment. Inability to properly account for these facets leaves such systems inaccurate in understanding of fine-grained proposition types. In this paper, we demonstrate an approach to annotating for four complex proposition types, namely normative claims, desires, future possibility, and reported speech. We develop a hybrid machine learning and human workflow for annotation that allows for efficient and reliable annotation of complex linguistic phenomena, and demonstrate with preliminary analysis of rhetorical strategies and structure in presidential debates. This new dataset and method can support technical researchers seeking more nuanced representations of argument, as well as argumentation theorists developing new quantitative analyses. Yohan Jo, Elijah Mayfield, Chris Reed 0001, Eduard H. Hovy |
LREC | 4 |
| 2020 | Forward and Backward Multimodal NMT for Improved Monolingual and Multilingual Cross-Modal RetrievalabstractWe explore methods to enrich the diversity of captions associated with pictures for learning improved visual-semantic embeddings (VSE) in cross-modal retrieval. In the spirit of "A picture is worth a thousand words", it would take dozens of sentences to parallel each picture's content adequately. But in fact, real-world multimodal datasets tend to provide only a few (typically, five) descriptions per image. For cross-modal retrieval, the resulting lack of diversity and coverage prevents systems from capturing the fine-grained inter-modal dependencies and intra-modal diversities in the shared VSE space. Using the fact that the encoder-decoder architectures in neural machine translation (NMT) have the capacity to enrich both monolingual and multilingual textual diversity, we propose a novel framework leveraging multimodal neural machine translation (MMT) to perform forward and backward translations based on salient visual objects to generate additional text-image pairs which enables training improved monolingual cross-modal retrieval (English-Image) and multilingual cross-modal retrieval (English-Image and German-Image) models. Experimental results show that the proposed framework can substantially and consistently improve the performance of state-of-the-art models on multiple datasets. The results also suggest that the models with multilingual VSE outperform the models with monolingual VSE. Po-Yao Huang 0001, Xiaojun Chang, Alex Hauptmann 0001, Eduard H. Hovy |
ICMR | 4 |
| 2020 | Unsupervised Data Augmentation for Consistency TrainingabstractSemi-supervised learning lately has shown much promise in improving deep learning models when labeled data is scarce. Common among recent approaches is the use of consistency training on a large amount of unlabeled data to constrain model predictions to be invariant to input noise. In this work, we present a new perspective on how to effectively noise unlabeled examples and argue that the quality of noising, specifically those produced by advanced data augmentation methods, plays a crucial role in semi-supervised learning. By substituting simple noising operations with advanced data augmentation methods such as RandAugment and back-translation, our method brings substantial improvements across six language and three vision tasks under the same consistency training framework. On the IMDb text classification dataset, with only 20 labeled examples, our method achieves an error rate of 4.20, outperforming the state-of-the-art model trained on 25,000 labeled examples. On a standard semi-supervised learning benchmark, CIFAR-10, our method outperforms all previous approaches and achieves an error rate of 5.43 with only 250 examples. Our method also combines well with transfer learning, e.g., when finetuning from BERT, and yields improvements in high-data regime, such as ImageNet, whether when there is only 10% labeled data or when a full labeled set with 1.3M extra unlabeled examples is used. Code is available at https://github.com/google-research/uda. Qizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong, Quoc V. Le |
NeurIPS | 3 |
| 2020 | Nested Named Entity Recognition via Second-best Sequence Learning and DecodingabstractWhen an entity name contains other names within it, the identification of all combinations of names can become difficult and expensive. We propose a new method to recognize not only outermost named entities but also inner nested ones. We design an objective function for training a neural model that treats the tag sequence for nested entities as the second best path within the span of their parent entity. In addition, we provide the decoding method for inference that extracts entities iteratively from outermost ones to inner ones in an outside-to-inside way. Our method has no additional hyperparameters to the conditional random field based model widely used for flat named entity recognition tasks. Experiments demonstrate that our method performs better than or at least as well as existing methods capable of handling nested entities, achieving F1-scores of 85.82%, 84.34%, and 77.36% on ACE-2004, ACE-2005, and GENIA datasets, respectively. Takashi Shibuya 0001, Eduard H. Hovy |
Trans. Assoc. Comput. Linguistics | 2 |
| 2020 | The role of knowledge in determining identity of long-tail entities
Filip Ilievski, Eduard H. Hovy, Piek Vossen, Stefan Schlobach, Qizhe Xie |
J. Web Semant. | 2 |
| 2019 | Neural Machine Translation with Adequacy-Oriented LearningabstractAlthough Neural Machine Translation (NMT) models have advanced state-of-the-art performance in machine translation, they face problems like the inadequate translation. We attribute this to that the standard Maximum Likelihood Estimation (MLE) cannot judge the real translation quality due to its several limitations. In this work, we propose an adequacyoriented learning mechanism for NMT by casting translation as a stochastic policy in Reinforcement Learning (RL), where the reward is estimated by explicitly measuring translation adequacy. Benefiting from the sequence-level training of RL strategy and a more accurate reward designed specifically for translation, our model outperforms multiple strong baselines, including (1) standard and coverage-augmented attention models with MLE-based training, and (2) advanced reinforcement and adversarial training strategies with rewards based on both word-level BLEU and character-level CHRF3. Quantitative and qualitative analyses on different language pairs and NMT architectures demonstrate the effectiveness and universality of the proposed approach. Xiang Kong, Zhaopeng Tu, Shuming Shi 0001, Eduard H. Hovy, Tong Zhang 0001 |
AAAI | 4 |
| 2019 | Fast and Simple Mixture of Softmaxes with BPE and Hybrid-LightRNN for Language GenerationabstractMixture of Softmaxes (MoS) has been shown to be effective at addressing the expressiveness limitation of Softmax-based models. Despite the known advantage, MoS is practically sealed by its large consumption of memory and computational time due to the need of computing multiple Softmaxes. In this work, we set out to unleash the power of MoS in practical applications by investigating improved word coding schemes, which could effectively reduce the vocabulary size and hence relieve the memory and computation burden. We show both BPE and our proposed Hybrid-LightRNN lead to improved encoding mechanisms that can halve the time and memory consumption of MoS without performance losses. With MoS, we achieve an improvement of 1.5 BLEU scores on IWSLT 2014 German-to-English corpus and an improvement of 0.76 CIDEr score on image captioning. Moreover, on the larger WMT 2014 machine translation dataset, our MoSboosted Transformer yields 29.6 BLEU score for English-toGerman and 42.1 BLEU score for English-to-French, outperforming the single-Softmax Transformer by 0.9 and 0.4 BLEU scores respectively and achieving the state-of-the-art result on WMT 2014 English-to-German task. Xiang Kong, Qizhe Xie, Zihang Dai, Eduard H. Hovy |
AAAI | 4 |
| 2019 | Women's Syntactic Resilience and Men's Grammatical Luck: Gender-Bias in Part-of-Speech Tagging and Dependency ParsingabstractSeveral linguistic studies have shown the prevalence of various lexical and grammatical patterns in texts authored by a person of a particular gender, but models for part-of-speech tagging and dependency parsing have still not adapted to account for these differences.To address this, we annotate the Wall Street Journal part of the Penn Treebank with the gender information of the articles' authors, and build taggers and parsers trained on this data that show performance differences in text written by men and women.Further analyses reveal numerous part-of-speech tags and syntactic relations whose prediction performances benefit from the prevalence of a specific gender in the training data.The results underscore the importance of accounting for gendered differences in syntactic tasks, and outline future venues for developing more accurate taggers and parsers.We release our data to the research community. Aparna Garimella, Carmen Banea, Eduard H. Hovy, Rada Mihalcea |
ACL (1) | 3 |
| 2019 | Exploring Numeracy in Word EmbeddingsabstractWord embeddings are now pervasive across NLP subfields as the de-facto method of forming text representataions.In this work, we show that existing embedding models are inadequate at constructing representations that capture salient aspects of mathematical meaning for numbers, which is important for language understanding.Numbers are ubiquitous and frequently appear in text.Inspired by cognitive studies on how humans perceive numbers, we develop an analysis framework to test how well word embeddings capture two essential properties of numbers: magnitude (e.g.3<4) and numeration (e.g.3=three).Our experiments reveal that most models capture an approximate notion of magnitude, but are inadequate at capturing numeration.We hope that our observations provide a starting point for the development of methods which better capture numeracy in NLP systems. Aakanksha Naik, Abhilasha Ravichander, Carolyn P. Rosé, Eduard H. Hovy |
ACL (1) | 4 |
| 2019 | Toward Comprehensive Understanding of a Sentiment Based on Human MotivesabstractIn sentiment detection, the natural language processing community has focused on determining holders, facets, and valences, but has paid little attention to the reasons for sentiment decisions. Our work considers human motives as the driver for human sentiments and addresses the problem of motive detection as the first step. Following a study in psychology, we define six basic motives that cover a wide range of topics appearing in review texts, annotate 1,600 texts in restaurant and laptop domains with the motives, and report the performance of baseline methods on this new dataset. We also show that cross-domain transfer learning boosts detection performance, which indicates that these universal motives exist across different domains. Naoki Otani, Eduard H. Hovy |
ACL (1) | 2 |
| 2019 | An Empirical Investigation of Structured Output Modeling for Graph-based Neural Dependency ParsingabstractIn this paper, we investigate the aspect of structured output modeling for the state-ofthe-art graph-based neural dependency parser (Dozat and Manning, 2017).With evaluations on 14 treebanks, we empirically show that global output-structured models can generally obtain better performance, especially on the metric of sentence-level Complete Match.However, probably because neural models already learn good global views of the inputs, the improvement brought by structured output modeling is modest. Zhisong Zhang, Xuezhe Ma, Eduard H. Hovy |
ACL (1) | 3 |
| 2019 | EQUATE: A Benchmark Evaluation Framework for Quantitative Reasoning in Natural Language InferenceabstractQuantitative reasoning is a higher-order reasoning skill that any intelligent natural language understanding system can reasonably be expected to handle.We present EQUATE 1 (Evaluating Quantitative Understanding Aptitude in Textual Entailment), a new framework for quantitative reasoning in textual entailment.We benchmark the performance of 9 published NLI models on EQUATE, and find that on average, state-of-the-art methods do not achieve an absolute improvement over a majority-class baseline, suggesting that they do not implicitly learn to reason with quantities.We establish a new baseline Q-REAS that manipulates quantities symbolically.In comparison to the best performing NLI model, it achieves success on numerical reasoning tests (+24.2%),but has limited verbal reasoning capabilities (-8.1%).We hope our evaluation framework will support the development of models of quantitative reasoning in language understanding. Abhilasha Ravichander, Aakanksha Naik, Carolyn P. Rosé, Eduard H. Hovy |
CoNLL | 4 |
| 2019 | Earlier Isn't Always Better: Sub-aspect Analysis on Corpus and System Biases in SummarizationabstractTaehee Jung, Dongyeop Kang, Lucas Mentch, Eduard Hovy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Taehee Jung, Dongyeop Kang, Lucas K. Mentch, Eduard H. Hovy |
EMNLP/IJCNLP (1) | 4 |
| 2019 | (Male, Bachelor) and (Female, Ph.D) have different connotations: Parallelly Annotated Stylistic Language Dataset with Multiple PersonasabstractDongyeop Kang, Varun Gangal, Eduard Hovy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dongyeop Kang, Varun Gangal, Eduard H. Hovy |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Linguistic Versus Latent Relations for Modeling Coherent Flow in ParagraphsabstractDongyeop Kang, Eduard Hovy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dongyeop Kang, Eduard H. Hovy |
EMNLP/IJCNLP (1) | 2 |
| 2019 | FlowSeq: Non-Autoregressive Conditional Sequence Generation with Generative FlowabstractXuezhe Ma, Chunting Zhou, Xian Li, Graham Neubig, Eduard Hovy. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xuezhe Ma, Chunting Zhou, Xian Li 0003, Graham Neubig, Eduard H. Hovy |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Learning Disentangled Representation in Latent Stochastic Models: A Case Study with Image CaptioningabstractMultimodal tasks require learning joint representation across modalities. In this paper, we present an approach to employ latent stochastic models for a multimodal task image captioning. Encoder Decoder models with stochastic latent variables are often faced with optimization issues such as latent collapse preventing them from realizing their full potential of rich representation learning and disentanglement. We present an approach to train such models by incorporating joint continuous and discrete representation in the prior distribution. We evaluate the performance of proposed approach on a multitude of metrics against vanilla latent stochastic models. We also perform a qualitative assessment and observe that the proposed approach indeed has the potential to learn composite information and explain novel combinations not seen in the training data. Nidhi Vyas, Sai Krishna Rallabandi, Lalitesh Morishetti, Eduard H. Hovy, Alan W. Black |
ICASSP | 4 |
| 2019 | MAE: Mutual Posterior-Divergence Regularization for Variational AutoEncoders
Xuezhe Ma, Chunting Zhou, Eduard H. Hovy |
ICLR (Poster) | 3 |
| 2019 | MaCow: Masked Convolutional Generative FlowabstractFlow-based generative models, conceptually attractive due to tractability of both the exact log-likelihood computation and latent-variable inference, and efficiency of both training and sampling, has led to a number of impressive empirical successes and spawned many advanced variants and theoretical investigations. Despite their computational efficiency, the density estimation performance of flow-based generative models significantly falls behind those of state-of-the-art autoregressive models. In this work, we introduce masked convolutional generative flow (MaCow), a simple yet effective architecture of generative flow using masked convolution. By restricting the local connectivity in a small kernel, MaCow enjoys the properties of fast and stable training, and efficient sampling, while achieving significant improvements over Glow for density estimation on standard image benchmarks, considerably narrowing the gap to autoregressive models. Xuezhe Ma, Xiang Kong, Shanghang Zhang, Eduard H. Hovy |
NeurIPS | 4 |
| 2019 | Discourse in Multimedia: A Case Study in Extracting Geometry Knowledge from TextbooksabstractTo ensure readability, text is often written and presented with due formatting. These text formatting devices help the writer to effectively convey the narrative. At the same time, these help the readers pick up the structure of the discourse and comprehend the conveyed information. There have been a number of linguistic theories on discourse structure of text. However, these theories only consider unformatted text. Multimedia text contains rich formatting features that can be leveraged for various NLP tasks. In this article, we study some of these discourse features in multimedia text and what communicative function they fulfill in the context. As a case study, we use these features to harvest structured subject knowledge of geometry from textbooks. We conclude that the discourse and text layout features provide information that is complementary to lexical semantic information. Finally, we show that the harvested structured knowledge can be used to improve an existing solver for geometry problems, making it more accurate as well as more explainable. Mrinmaya Sachan, Avinava Dubey, Eduard H. Hovy, Tom M. Mitchell, Dan Roth 0001, Eric P. Xing |
Comput. Linguistics | 3 |
| 2018 | SPINE: SParse Interpretable Neural EmbeddingsabstractPrediction without justification has limited utility. Much of the success of neural models can be attributed to their ability to learn rich, dense and expressive representations. While these representations capture the underlying complexity and latent trends in the data, they are far from being interpretable. We propose a novel variant of denoising k-sparse autoencoders that generates highly efficient and interpretable distributed word representations (word embeddings), beginning with existing word representations from state-of-the-art methods like GloVe and word2vec. Through large scale human evaluation, we report that our resulting word embedddings are much more interpretable than the original GloVe and word2vec embeddings. Moreover, our embeddings outperform existing popular word embeddings on a diverse suite of benchmark downstream tasks. Anant Subramanian, Danish Pruthi, Harsh Jhamtani, Taylor Berg-Kirkpatrick, Eduard H. Hovy |
AAAI | 5 |
| 2018 | AdvEntuRe: Adversarial Training for Textual Entailment with Knowledge-Guided ExamplesabstractWe consider the problem of learning textual entailment models with limited supervision (5K-10K training examples), and present two complementary approaches for it.First, we propose knowledge-guided adversarial example generators for incorporating large lexical resources in entailment models via only a handful of rule templates.Second, to make the entailment model-a discriminator-more robust, we propose the first GAN-style approach for training it using a natural language example generator that iteratively adjusts based on the discriminator's performance.We demonstrate effectiveness using two entailment datasets, where the proposed methods increase accuracy by 4.7% on SciTail and by 2.8% on a 1% training sub-sample of SNLI.Notably, even a single hand-written rule, negate, improves the accuracy on the negation examples in SNLI by 6.1%.P: The dog did not eat all of the chickens.H: The dog ate all of the chickens.S: entails (score 56:5%) P: The red box is in the blue box.H: The blue box is in the red box. Dongyeop Kang, Tushar Khot, Ashish Sabharwal, Eduard H. Hovy |
ACL (1) | 4 |
| 2018 | Stack-Pointer Networks for Dependency ParsingabstractWe introduce a novel architecture for dependency parsing: stack-pointer networks (STACKPTR).Combining pointer networks (Vinyals et al., 2015) with an internal stack, the proposed model first reads and encodes the whole sentence, then builds the dependency tree top-down (from root-to-leaf) in a depth-first fashion.The stack tracks the status of the depthfirst search and the pointer networks select one child for the word at the top of the stack at each step.The STACKPTR parser benefits from the information of the whole sentence and all previously derived subtree structures, and removes the leftto-right restriction in classical transitionbased parsers.Yet, the number of steps for building any (including non-projective) parse tree is linear in the length of the sentence just as other transition-based parsers, yielding an efficient decoding algorithm with O(n 2 ) time complexity.We evaluate our model on 29 treebanks spanning 20 languages and different dependency annotation schemas, and achieve state-of-theart performance on 21 of them. Xuezhe Ma, Zecong Hu, Jingzhou Liu, Nanyun Peng 0001, Graham Neubig, Eduard H. Hovy |
ACL (1) | 6 |
| 2018 | Learning to Generate Move-by-Move Commentary for Chess Games from Large-Scale Social Forum DataabstractHarsh Jhamtani, Varun Gangal, Eduard Hovy, Graham Neubig, Taylor Berg-Kirkpatrick. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Harsh Jhamtani, Varun Gangal, Eduard H. Hovy, Graham Neubig, Taylor Berg-Kirkpatrick |
ACL (1) | 3 |
| 2018 | From Credit Assignment to Entropy Regularization: Two New Algorithms for Neural Sequence PredictionabstractIn this work, we study the credit assignment problem in reward augmented maximum likelihood (RAML) learning, and establish a theoretical equivalence between the token-level counterpart of RAML and the entropy regularized reinforcement learning.Inspired by the connection, we propose two sequence prediction algorithms, one extending RAML with fine-grained credit assignment and the other improving Actor-Critic with a systematic entropy regularization.On two benchmark datasets, we show the proposed algorithms outperform RAML and Actor-Critic respectively, providing new alternatives to sequence prediction. Zihang Dai, Qizhe Xie, Eduard H. Hovy |
ACL (1) | 3 |
| 2018 | Graph Based Decoding for Event Sequencing and Coreference ResolutionabstractEvents in text documents are interrelated in complex ways. In this paper, we study two types of relation: Event Coreference and Event Sequencing. We show that the popular tree-like decoding structure for automated Event Coreference is not suitable for Event Sequencing. To this end, we propose a graph-based decoding algorithm that is applicable to both tasks. The new decoding algorithm supports flexible feature sets for both tasks. Empirically, our event coreference system has achieved state-of-the-art performance on the TAC-KBP 2015 event coreference task and our event sequencing system beats a strong temporal-based, oracle-informed baseline. We discuss the challenges of studying these event relations. Zhengzhong Liu 0001, Teruko Mitamura, Eduard H. Hovy |
COLING | 3 |
| 2018 | Low-resource Cross-lingual Event Type Detection via Distant Supervision with Minimal EffortabstractThe use of machine learning for NLP generally requires resources for training. Tasks performed in a low-resource language usually rely on labeled data in another, typically resource-rich, language. However, there might not be enough labeled data even in a resource-rich language such as English. In such cases, one approach is to use a hand-crafted approach that utilizes only a small bilingual dictionary with minimal manual verification to create distantly supervised data. Another is to explore typical machine learning techniques, for example adversarial training of bilingual word representations. We find that in event-type detection task—the task to classify [parts of] documents into a fixed set of labels—they give about the same performance. We explore ways in which the two methods can be complementary and also see how to best utilize a limited budget for manual annotation to maximize performance gain. Aldrian Obaja Muis, Naoki Otani, Nidhi Vyas, Ruochen Xu, Yiming Yang 0002, Teruko Mitamura, Eduard H. Hovy |
COLING | 7 |
| 2018 | Automatic Event Salience IdentificationabstractIdentifying the salience (i.e.importance) of discourse units is an important task in language understanding.While events play important roles in text documents, little research exists on analyzing their saliency status.This paper empirically studies the Event Salience task and proposes two salience detection models based on content similarities and discourse relations.The first is a feature based salience model that incorporates similarities among discourse units.The second is a neural model that captures more complex relations between discourse units.Tested on our new largescale event salience corpus, both methods significantly outperform the strong frequency baseline, while our neural model further improves the feature based one by a large margin.Our analyses demonstrate that our neural model captures interesting connections between salience and discourse unit relations (e.g., scripts and frame structures). Zhengzhong Liu 0001, Chenyan Xiong, Teruko Mitamura, Eduard H. Hovy |
EMNLP | 4 |
| 2018 | Large-scale Cloze Test Dataset Created by TeachersabstractCloze tests are widely adopted in language exams to evaluate students' language proficiency.In this paper, we propose the first large-scale human-created cloze test dataset CLOTH 1 2 , containing questions used in middle-school and high-school language exams.With missing blanks carefully created by teachers and candidate choices purposely designed to be nuanced, CLOTH requires a deeper language understanding and a wider attention span than previously automaticallygenerated cloze datasets.We test the performance of dedicatedly designed baseline models including a language model trained on the One Billion Word Corpus and show humans outperform them by a significant margin.We investigate the source of the performance gap, trace model deficiencies to some distinct properties of CLOTH, and identify the limited ability of comprehending the long-term context to be the key bottleneck. Qizhe Xie, Guokun Lai, Zihang Dai, Eduard H. Hovy |
EMNLP | 4 |
| 2018 | Reading Agents that Hunger for Knowledge
Eduard H. Hovy |
ICAART (1) | 1 |
| 2018 | A Dataset of Peer Reviews (PeerRead): Collection, Insights and NLP ApplicationsabstractDongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard Hovy, Roy Schwartz. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard H. Hovy, Roy Schwartz 0001 |
NAACL-HLT | 6 |
| 2018 | The ARIEL-CMU situation frame detection pipeline for LoReHLT16: a model translation approach
Patrick Littell, Ruochen Xu, Zaid Sheikh, David R. Mortensen, Lori S. Levin, Francis M. Tyers, Hiroaki Hayashi, Graham Horwood, Steve Sloto, Emily Tagtow, Alan W. Black, Yiming Yang 0002, Teruko Mitamura, Eduard H. Hovy |
Mach. Transl. | 15 |
| 2017 | Ontology-Aware Token Embeddings for Prepositional Phrase AttachmentabstractType-level word embeddings use the same set of parameters to represent all instances of a word regardless of its context, ignoring the inherent lexical ambiguity in language.Instead, we embed semantic concepts (or synsets) as defined in WordNet and represent a word token in a particular context by estimating a distribution over relevant semantic concepts.We use the new, context-sensitive embeddings in a model for predicting prepositional phrase (PP) attachments and jointly learn the concept embeddings and model parameters.We show that using context-sensitive embeddings improves the accuracy of the PP attachment model by 5.4% absolute points, which amounts to a 34.4% relative reduction in errors. Pradeep Dasigi, Waleed Ammar, Chris Dyer, Eduard H. Hovy |
ACL (1) | 4 |
| 2017 | An Interpretable Knowledge Transfer Model for Knowledge Base CompletionabstractKnowledge bases are important resources for a variety of natural language processing tasks but suffer from incompleteness.We propose a novel embedding model, ITransF, to perform knowledge base completion.Equipped with a sparse attention mechanism, ITransF discovers hidden concepts of relations and transfer statistical strength through the sharing of concepts.Moreover, the learned associations between relations and concepts, which are represented by sparse attention vectors, can be interpreted easily.We evaluate ITransF on two benchmark datasets-WN18 and FB15k for knowledge base completion and obtains improvements on both the mean rank and Hits@10 metrics, over all baselines that do not use additional information. Qizhe Xie, Xuezhe Ma, Zihang Dai, Eduard H. Hovy |
ACL (1) | 4 |
| 2017 | JointSem: Combining Query Entity Linking and Entity based Document RankingabstractEntity-based ranking systems often employ entity linking systems to align entities to query and documents. Previously, entity linking systems were not designed specifically for search engines and were mostly used as a preprocessing step. This work presents JointSem, a joint semantic ranking system that combines query entity linking and entity-based document ranking. In JointSem, the spotting and linking signals are used to describe the importance of candidate entities in the query, and the linked entities are utilized to provide additional ranking features for the documents. The linking signals and the ranking signals are combined by a joint learning-to-rank model, and the whole system is fully optimized towards end-to-end ranking performance. Experiments on TREC Web Track datasets demonstrate the effectiveness of joint learning of entity linking and entity-based ranking. Chenyan Xiong, Zhengzhong Liu 0001, Jamie Callan, Eduard H. Hovy |
CIKM | 4 |
| 2017 | Charmanteau: Character Embedding Models For Portmanteau CreationabstractPortmanteaus are a word formation phenomenon where two words are combined to form a new word.We propose character-level neural sequence-tosequence (S2S) methods for the task of portmanteau generation that are end-toend-trainable, language independent, and do not explicitly use additional phonetic information.We propose a noisy-channelstyle model, which allows for the incorporation of unsupervised word lists, improving performance over a standard sourceto-target model.This model is made possible by an exhaustive candidate generation strategy specifically enabled by the features of the portmanteau task.Experiments find our approach superior to a state-of-the-art FST-based baseline with respect to ground truth accuracy and human evaluation. Varun Gangal, Harsh Jhamtani, Graham Neubig, Eduard H. Hovy, Eric Nyberg |
EMNLP | 4 |
| 2017 | Detecting and Explaining Causes From Text For a Time Series EventabstractExplaining underlying causes or effects about events is a challenging but valuable task.We define a novel problem of generating explanations of a time series event by (1) searching cause and effect relationships of the time series with textual data and (2) constructing a connecting chain between them to generate an explanation.To detect causal features from text, we propose a novel method based on the Granger causality of time series between features extracted from text such as N-grams, topics, sentiments, and their composition.The generation of the sequence of causal entities requires a commonsense causative knowledge base with efficient reasoning.To ensure good interpretability and appropriate lexical usage we combine symbolic and neural representations, using a neural reasoning algorithm trained on commonsense causal tuples to predict the next cause step.Our quantitative and human analysis show empirical evidence that our method successfully extracts meaningful causality relationships between time series with textual features and generates appropriate explanation between them. Dongyeop Kang, Varun Gangal, Ang Lu, Eduard H. Hovy |
EMNLP | 5 |
| 2017 | RACE: Large-scale ReAding Comprehension Dataset From ExaminationsabstractWe present RACE, a new dataset for benchmark evaluation of methods in the reading comprehension task.Collected from the English exams for middle and high school Chinese students in the age range between 12 to 18, RACE consists of near 28,000 passages and near 100,000 questions generated by human experts (English instructors), and covers a variety of topics which are carefully designed for evaluating the students' ability in understanding and reasoning.In particular, the proportion of questions that requires reasoning is much larger in RACE than that in other benchmark datasets for reading comprehension, and there is a significant gap between the performance of the state-of-the-art models (43%) and the ceiling human performance (95%).We hope this new dataset can serve as a valuable resource for research and evaluation in machine comprehension. Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang 0002, Eduard H. Hovy |
EMNLP | 5 |
| 2017 | Identifying Semantic Edit Intentions from Revisions in WikipediaabstractMost studies on human editing focus merely on syntactic revision operations, failing to capture the intentions behind revision changes, which are essential for facilitating the single and collaborative writing process.In this work, we develop in collaboration with Wikipedia editors a 13-category taxonomy of the semantic intention behind edits in Wikipedia articles.Using labeled article edits, we build a computational classifier of intentions that achieved a micro-averaged F1 score of 0.621.We use this model to investigate edit intention effectiveness: how different types of edits predict the retention of newcomers and changes in the quality of articles, two key concerns for Wikipedia today.Our analysis shows that the types of edits that users make in their first session predict their subsequent survival as Wikipedia editors, and articles in different stages need different types of edits. Diyi Yang, Aaron Halfaker, Robert E. Kraut, Eduard H. Hovy |
EMNLP | 4 |
| 2017 | Calibrating Energy-based Generative Adversarial Networks
Zihang Dai, Amjad Almahairi, Philip Bachman, Eduard H. Hovy, Aaron C. Courville |
ICLR (Poster) | 4 |
| 2017 | Dropout with Expectation-linear Regularization
Xuezhe Ma, Yingkai Gao, Zhiting Hu, Yaoliang Yu, Yuntian Deng, Eduard H. Hovy |
ICLR (Poster) | 6 |
| 2017 | Neural Probabilistic Model for Non-projective MST ParsingabstractIn this paper, we propose a probabilistic parsing model that defines a proper conditional probability distribution over non-projective dependency trees for a given sentence, using neural representations as inputs. The neural network architecture is based on bi-directional LSTMCNNs, which automatically benefits from both word- and character-level representations, by using a combination of bidirectional LSTMs and CNNs. On top of the neural network, we introduce a probabilistic structured layer, defining a conditional log-linear model over non-projective trees. By exploiting Kirchhoff’s Matrix-Tree Theorem (Tutte, 1984), the partition functions and marginals can be computed efficiently, leading to a straightforward end-to-end model training procedure via back-propagation. We evaluate our model on 17 different datasets, across 14 different languages. Our parser achieves state-of-the-art parsing performance on nine datasets. Xuezhe Ma, Eduard H. Hovy |
IJCNLP(1) | 2 |
| 2017 | Controllable Invariance through Adversarial Feature LearningabstractLearning meaningful representations that maintain the content necessary for a particular task while filtering away detrimental variations is a problem of great interest in machine learning. In this paper, we tackle the problem of learning representations invariant to a specific factor or trait of data. The representation learning process is formulated as an adversarial minimax game. We analyze the optimal equilibrium of such a game and find that it amounts to maximizing the uncertainty of inferring the detrimental factor given the representation while maximizing the certainty of making task-specific predictions. On three benchmark tasks, namely fair and bias-free classification, language-independent generation, and lighting-independent image classification, we show that the proposed framework induces an invariant representation, and leads to better generalization evidenced by the improved performance. Qizhe Xie, Zihang Dai, Yulun Du, Eduard H. Hovy, Graham Neubig |
NIPS | 4 |
| 2017 | Finding Structure in Figurative Language: Metaphor Detection with Topic-based FramesabstractIn this paper, we present a novel and highly effective method for induction and application of metaphor frame templates as a step toward detecting metaphor in extended discourse.We infer implicit facets of a given metaphor frame using a semisupervised bootstrapping approach on an unlabeled corpus.Our model applies this frame facet information to metaphor detection, and achieves the state-of-the-art performance on a social media dataset when building upon other proven features in a nonlinear machine learning model.In addition, we illustrate the mechanism through which the frame and topic information enable the more accurate metaphor detection. Hyeju Jang, Korte Maki, Eduard H. Hovy, Carolyn P. Rosé |
SIGDIAL Conference | 3 |
| 2016 | Harnessing Deep Neural Networks with Logic RulesabstractCombining deep neural networks with structured logic rules is desirable to harness flexibility and reduce uninterpretability of the neural models.We propose a general framework capable of enhancing various types of neural networks (e.g., CNNs and RNNs) with declarative first-order logic rules.Specifically, we develop an iterative distillation method that transfers the structured information of logic rules into the weights of neural networks.We deploy the framework on a CNN for sentiment analysis, and an RNN for named entity recognition.With a few highly intuitive rules, we obtain substantial improvements and achieve state-of-the-art or comparable results to previous best-performing systems. Zhiting Hu, Xuezhe Ma, Zhengzhong Liu 0001, Eduard H. Hovy, Eric P. Xing |
ACL (1) | 4 |
| 2016 | Tables as Semi-structured Knowledge for Question AnsweringabstractQuestion answering requires access to a knowledge base to check facts and reason about information.Knowledge in the form of natural language text is easy to acquire, but difficult for automated reasoning.Highly-structured knowledge bases can facilitate reasoning, but are difficult to acquire.In this paper we explore tables as a semi-structured formalism that provides a balanced compromise to this tradeoff.We first use the structure of tables to guide the construction of a dataset of over 9000 multiple-choice questions with rich alignment annotations, easily and efficiently via crowd-sourcing.We then use this annotated data to train a semistructured feature-driven model for question answering that uses tables as a knowledge base.In benchmark evaluations, we significantly outperform both a strong unstructured retrieval baseline and a highlystructured Markov Logic Network model. Sujay Kumar Jauhar, Peter D. Turney, Eduard H. Hovy |
ACL (1) | 3 |
| 2016 | End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRFabstractState-of-the-art sequence labeling systems traditionally require large amounts of taskspecific knowledge in the form of handcrafted features and data pre-processing.In this paper, we introduce a novel neutral network architecture that benefits from both word-and character-level representations automatically, by using combination of bidirectional LSTM, CNN and CRF.Our system is truly end-to-end, requiring no feature engineering or data preprocessing, thus making it applicable to a wide range of sequence labeling tasks.We evaluate our system on two data sets for two sequence labeling tasks -Penn Treebank WSJ corpus for part-of-speech (POS) tagging and CoNLL 2003 corpus for named entity recognition (NER).We obtain state-of-the-art performance on both datasets -97.55% accuracy for POS tagging and 91.21% F1 for NER. Xuezhe Ma, Eduard H. Hovy |
ACL (1) | 2 |
| 2016 | The Creation and Analysis of a Website Privacy Policy CorpusabstractShomir Wilson, Florian Schaub, Aswarth Abhilash Dara, Frederick Liu, Sushain Cherivirala, Pedro Giovanni Leon, Mads Schaarup Andersen, Sebastian Zimmeck, Kanthashree Mysore Sathyendra, N. Cameron Russell, Thomas B. Norton, Eduard Hovy, Joel Reidenberg, Norman Sadeh. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Shomir Wilson, Florian Schaub, Aswarth Abhilash Dara, Frederick Liu, Sushain Cherivirala, Pedro Giovanni Leon, Mads Schaarup Andersen, Sebastian Zimmeck, Kanthashree Mysore Sathyendra, N. Cameron Russell, Thomas B. Norton, Eduard H. Hovy, Joel R. Reidenberg, Norman M. Sadeh |
ACL (1) | 12 |
| 2016 | Who Did What: Editor Role Identification in Wikipedia
Diyi Yang, Aaron Halfaker, Robert E. Kraut, Eduard H. Hovy |
ICWSM | 4 |
| 2016 | Edit Categories and Editor Role Identification in Wikipedia
Diyi Yang, Aaron Halfaker, Robert E. Kraut, Eduard H. Hovy |
LREC | 4 |
| 2016 | Visualizing and Understanding Neural Models in NLPabstractWhile neural networks have been successfully applied to many NLP tasks the resulting vectorbased models are very difficult to interpret.For example it's not clear how they achieve compositionality, building sentence meaning from the meanings of words and phrases.In this paper we describe strategies for visualizing compositionality in neural models for NLP, inspired by similar work in computer vision.We first plot unit values to visualize compositionality of negation, intensification, and concessive clauses, allowing us to see wellknown markedness asymmetries in negation.We then introduce methods for visualizing a unit's salience, the amount that it contributes to the final composed meaning from first-order derivatives.Our general-purpose methods may have wide applications for understanding compositionality and other semantic properties of deep networks. Jiwei Li 0001, Xinlei Chen, Eduard H. Hovy, Daniel Jurafsky |
HLT-NAACL | 3 |
| 2016 | Unsupervised Ranking Model for Entity Coreference ResolutionabstractCoreference resolution is one of the first stages in deep language understanding and its importance has been well recognized in the natural language processing community. In this paper, we propose a generative, unsupervised ranking model for entity coreference resolution by introducing resolution mode variables. Our unsupervised system achieves 58.44% F1 score of the CoNLL metric on the English data from the CoNLL-2012 shared task (Pradhan et al., 2012), outperforming the Stanford deterministic system (Lee et al., 2013) by 3.01%. Xuezhe Ma, Zhengzhong Liu 0001, Eduard H. Hovy |
HLT-NAACL | 3 |
| 2016 | Hierarchical Attention Networks for Document ClassificationabstractZichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, Eduard Hovy. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Diyi Yang, Chris Dyer, Xiaodong He 0001, Alexander J. Smola, Eduard H. Hovy |
HLT-NAACL | 6 |
| 2015 | When Are Tree Structures Necessary for Deep Learning of Representations?abstractRecursive neural models, which use syntactic parse trees to recursively generate representations bottom-up, are a popular architecture.However there have not been rigorous evaluations showing for exactly which tasks this syntax-based method is appropriate.In this paper, we benchmark recursive neural models against sequential recurrent neural models, enforcing applesto-apples comparison as much as possible.We investigate 4 tasks: (1) sentiment classification at the sentence level and phrase level; (2) matching questions to answerphrases; (3) discourse parsing; (4) semantic relation extraction.Our goal is to understand better when, and why, recursive models can outperform simpler models.We find that recursive models help mainly on tasks (like semantic relation extraction) that require longdistance connection modeling, particularly on very long sequences.We then introduce a method for allowing recurrent models to achieve similar performance: breaking long sentences into clause-like units at punctuation and processing them separately before combining.Our results thus help understand the limitations of both classes of models, and suggest directions for improving recurrent models. Jiwei Li 0001, Thang Luong, Daniel Jurafsky, Eduard H. Hovy |
EMNLP | 4 |
| 2015 | Efficient Inner-to-outer Greedy Algorithm for Higher-order Labeled Dependency ParsingabstractMany NLP systems use dependency parsers as critical components.Jonit learning parsers usually achieve better parsing accuracies than two-stage methods.However, classical joint parsing algorithms significantly increase computational complexity, which makes joint learning impractical.In this paper, we proposed an efficient dependency parsing algorithm that is capable of capturing multiple edge-label features, while maintaining low computational complexity.We evaluate our parser on 14 different languages.Our parser consistently obtains more accurate results than three baseline systems and three popular, off-the-shelf parsers. Xuezhe Ma, Eduard H. Hovy |
EMNLP | 2 |
| 2015 | Humor Recognition and Humor Anchor ExtractionabstractHumor is an essential component in personal communication. How to create computational models to discover the structures behind humor, recognize humor and even extract humor anchors remains a challenge. In this work, we first identify several semantic structures behind humor and design sets of features for each structure, and next employ a com-putational approach to recognize humor. Furthermore, we develop a simple and effective method to extract anchors that enable humor in a sentence. Experiments conducted on two datasets demonstrate that our humor recognizer is effective in automatically distinguishing between humorous and non-humorous texts and our extracted humor anchors correlate quite well with human annotations. 1 Diyi Yang, Alon Lavie, Chris Dyer, Eduard H. Hovy |
EMNLP | 4 |
| 2015 | An Active Learning Approach to Coreference Resolution
Mrinmaya Sachan, Eduard H. Hovy, Eric P. Xing |
IJCAI | 2 |
| 2015 | Retrofitting Word Vectors to Semantic LexiconsabstractManaal Faruqui, Jesse Dodge, Sujay Kumar Jauhar, Chris Dyer, Eduard Hovy, Noah A. Smith. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Manaal Faruqui, Jesse Dodge, Sujay Kumar Jauhar, Chris Dyer, Eduard H. Hovy, Noah A. Smith |
HLT-NAACL | 5 |
| 2015 | Ontologically Grounded Multi-sense Representation Learning for Semantic Vector Space ModelsabstractWords are polysemous. However, most approaches to representation learning for lexical semantics assign a single vector to every surface word type. Meanwhile, lexical ontologies such as WordNet provide a source of complementary knowledge to distributional information, including a word sense inventory. In this paper we propose two novel and general approaches for generating sense-specific word embeddings that are grounded in an ontology. The first applies graph smoothing as a postprocessing step to tease the vectors of different senses apart, and is applicable to any vector space model. The second adapts predictive maximum likelihood models that learn word embeddings with latent variables representing senses grounded in an specified ontology. Empirical results on lexical semantic tasks show that our approaches effectively captures information from both the ontology and distributional statistics. Moreover, in most cases our sense-specific models outperform other models we compare against. Sujay Kumar Jauhar, Chris Dyer, Eduard H. Hovy |
HLT-NAACL | 3 |
| 2014 | Towards a General Rule for Identifying Deceptive Opinion SpamabstractConsumers' purchase decisions are increasingly influenced by user-generated online reviews.Accordingly, there has been growing concern about the potential for posting deceptive opinion spamfictitious reviews that have been deliberately written to sound authentic, to deceive the reader.In this paper, we explore generalized approaches for identifying online deceptive opinion spam based on a new gold standard dataset, which is comprised of data from three different domains (i.e.Hotel, Restaurant, Doctor), each of which contains three types of reviews, i.e. customer generated truthful reviews, Turker generated deceptive reviews and employee (domain-expert) generated deceptive reviews.Our approach tries to capture the general difference of language usage between deceptive and truthful reviews, which we hope will help customers when making purchase decisions and review portal operators, such as TripAdvisor or Yelp, investigate possible fraudulent activity on their sites.1 Jiwei Li 0001, Myle Ott, Claire Cardie, Eduard H. Hovy |
ACL (1) | 4 |
| 2014 | Weakly Supervised User Profile Extraction from TwitterabstractWhile user attribute extraction on social media has received considerable attention, existing approaches, mostly supervised, encounter great difficulty in obtaining gold standard data and are therefore limited to predicting unary predicates (e.g., gender).In this paper, we present a weaklysupervised approach to user profile extraction from Twitter.Users' profiles from social media websites such as Facebook or Google Plus are used as a distant source of supervision for extraction of their attributes from user-generated text.In addition to traditional linguistic features used in distant supervision for information extraction, our approach also takes into account network information, a unique opportunity offered by social media.We test our algorithm on three attribute domains: spouse, education and job; experimental results demonstrate our approach is able to make accurate predictions for users' attributes based on their tweets.1• We experimentally demonstrate the effectiveness of our approach on 3 relations: SPOUSE, JOB and EDUCATION.The remainder of this paper is organized as follows: We summarize related work in Section 2. The creation of our dataset is described in Section 3. The details of our model are presented in Section 4. We present experimental results in Section 5 and conclude in Section 6. Jiwei Li 0001, Alan Ritter, Eduard H. Hovy |
ACL (1) | 3 |
| 2014 | Vector space semantics with frequency-driven motifsabstractTraditional models of distributional semantics suffer from computational issues such as data sparsity for individual lexemes and complexities of modeling semantic composition when dealing with structures larger than single lexical items.In this work, we present a frequencydriven paradigm for robust distributional semantics in terms of semantically cohesive lineal constituents, or motifs.The framework subsumes issues such as differential compositional as well as noncompositional behavior of phrasal consituents, and circumvents some problems of data sparsity by design.We design a segmentation model to optimally partition a sentence into lineal constituents, which can be used to define distributional contexts that are less noisy, semantically more interpretable, and linguistically disambiguated.Hellinger PCA embeddings learnt using the framework show competitive results on empirical tasks. Eduard H. Hovy |
ACL (1) | 2 |
| 2014 | What a Nasty Day: Exploring Mood-Weather Relationship from TwitterabstractWhile it has long been believed in psychology that weather somehow influences human's mood, the debates have been going on for decades about how they are correlated. In this paper, we try to study this long-lasting topic by harnessing a new source of data compared from traditional psychological researches: Twitter. We analyze 2 years' twitter data collected by twitter API which amounts to 10% of all postings and try to reveal the correlations between multiple dimensional structure of human mood with meteorological effects. Some of our findings confirm existing hypotheses, while others contradict them. We are hopeful that our approach, along with the new data source, can shed on the long-going debates on weather-mood correlation. Jiwei Li 0001, Eduard H. Hovy |
CIKM | 3 |
| 2014 | Modeling Newswire Events using Neural Networks for Anomaly Detection
Pradeep Dasigi, Eduard H. Hovy |
COLING | 2 |
| 2014 | Unsupervised Word Sense Induction using Distributional Statistics
Kartik Goyal, Eduard H. Hovy |
COLING | 2 |
| 2014 | Inducing Latent Semantic Relations for Structured Distributional Semantics
Sujay Kumar Jauhar, Eduard H. Hovy |
COLING | 2 |
| 2014 | Sentiment Analysis on the People's DailyabstractWe propose a semi-supervised bootstrap-ping algorithm for analyzing China’s for-eign relations from the People’s Daily. Our approach addresses sentiment tar-get clustering, subjective lexicons extrac-tion and sentiment prediction in a unified framework. Different from existing algo-rithms in the literature, time information is considered in our algorithm through a hierarchical bayesian model to guide the bootstrapping approach. We are hopeful that our approach can facilitate quantita-tive political analysis conducted by social scientists and politicians. 1 Jiwei Li 0001, Eduard H. Hovy |
EMNLP | 2 |
| 2014 | A Model of Coherence Based on Distributed Sentence RepresentationabstractCoherence is what makes a multi-sentence text meaningful, both logically and syntactically.To solve the challenge of ordering a set of sentences into coherent order, existing approaches focus mostly on defining and using sophisticated features to capture the cross-sentence argumentation logic and syntactic relationships.But both argumentation semantics and crosssentence syntax (such as coreference and tense rules) are very hard to formalize.In this paper, we introduce a neural network model for the coherence task based on distributed sentence representation.The proposed approach learns a syntacticosemantic representation for sentences automatically, using either recurrent or recursive neural networks.The architecture obviated the need for feature engineering, and learns sentence representations, which are to some extent able to capture the 'rules' governing coherent sentence structure.The proposed approach outperforms existing baselines and generates the stateof-art performance in standard coherence evaluation tasks 1 . Jiwei Li 0001, Eduard H. Hovy |
EMNLP | 2 |
| 2014 | Recursive Deep Models for Discourse ParsingabstractText-level discourse parsing remains a challenge: most approaches employ fea-tures that fail to capture the intentional, se-mantic, and syntactic aspects that govern discourse coherence. In this paper, we pro-pose a recursive model for discourse pars-ing that jointly models distributed repre-sentations for clauses, sentences, and en-tire discourses. The learned representa-tions can to some extent learn the seman-tic and intentional import of words and larger discourse units automatically,. The proposed framework obtains comparable performance regarding standard discours-ing parsing evaluations when compared against current state-of-art systems. 1 Jiwei Li 0001, Rumeng Li, Eduard H. Hovy |
EMNLP | 3 |
| 2014 | Major Life Event Extraction from Twitter based on Congratulations/Condolences Speech ActsabstractSocial media websites provide a platform for anyone to describe significant events taking place in their lives in realtime.Currently, the majority of personal news and life events are published in a textual format, motivating information extraction systems that can provide a structured representations of major life events (weddings, graduation, etc. . .).This paper demonstrates the feasibility of accurately extracting major life events.Our system extracts a fine-grained description of users' life events based on their published tweets.We are optimistic that our system can help Twitter users more easily grasp information from users they take interest in following and also facilitate many downstream applications, for example realtime friend recommendation. Jiwei Li 0001, Alan Ritter, Claire Cardie, Eduard H. Hovy |
EMNLP | 4 |
| 2014 | Detecting Subevent Structure for Event Coreference Resolution
Jun Araki, Zhengzhong Liu 0001, Eduard H. Hovy, Teruko Mitamura |
LREC | 3 |
| 2014 | A Corpus of Participant Roles in Contentious Discussions
Archna Bhatia, Angelique Rein, Eduard H. Hovy |
LREC | 4 |
| 2014 | Supervised Within-Document Event Coreference using Information Propagation
Zhengzhong Liu 0001, Jun Araki, Eduard H. Hovy, Teruko Mitamura |
LREC | 3 |
| 2014 | Spatial compactness meets topical consistency: jointly modeling links and content for community detectionabstractIn this paper, we address the problem of discovering topically meaningful, yet compact (densely connected) communities in a social network. Assuming the social network to be an integer-weighted graph (where the weights can be intuitively defined as the number of common friends, followers, documents exchanged, etc.), we transform the social network to a more efficient representation. In this new representation, each user is a bag of her one-hop neighbors. We propose a mixed-membership model to identify compact communities using this transformation. Next, we augment the representation and the model to incorporate user-content information imposing topical consistency in the communities. In our model a user can belong to multiple communities and a community can participate in multiple topics. This allows us to discover community memberships as well as community and user interests. Our method outperforms other well known baselines on two real-world social networks. Finally, we also provide a fast, parallel approximation of the same. Mrinmaya Sachan, Avinava Dubey, Eric P. Xing, Eduard H. Hovy |
WSDM | 5 |
| 2013 | Automatic Interpretation of the English Possessive
Stephen Tratz, Eduard H. Hovy |
ACL (1) | 2 |
| 2013 | A Walk-Based Semantically Enriched Tree Kernel Over Distributed Word RepresentationsabstractIn this paper, we propose a walk-based graph kernel that generalizes the notion of treekernels to continuous spaces.Our proposed approach subsumes a general framework for word-similarity, and in particular, provides a flexible way to incorporate distributed representations.Using vector representations, such an approach captures both distributional semantic similarities among words as well as the structural relations between them (encoded as the structure of the parse tree).We show an efficient formulation to compute this kernel using simple matrix operations.We present our results on three diverse NLP tasks, showing state-of-the-art results. Dirk Hovy, Eduard H. Hovy |
EMNLP | 3 |
| 2013 | Analysis and modeling of "focus" in contextabstractThis paper uses a crowd-sourced definition of a speech phe-nomenon we have called “focus”. Given sentences, text and speech, in isolation and in context, we asked annotators to iden-tify what we term the “focus ” word. We present their consis-tency in identifying the focused word, when presented with text or speech stimuli. We then build models to show how well we predict that focus word from lexical (and higher) level features. Also, using spectral and prosodic information, we show the dif-ferences in these focus words when spoken with and without context. Finally, we show how we can improve speech synthe-sis of these utterances given focus information. Dirk Hovy, Gopala Krishna Anumanchipalli, Alok Parlikar, Caroline Vaughn, Adam C. Lammert, Eduard H. Hovy, Alan W. Black |
INTERSPEECH | 6 |
| 2013 | Learning Whom to Trust with MACE
Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, Eduard H. Hovy |
HLT-NAACL | 4 |
| 2013 | Editorial
Eduard H. Hovy, Roberto Navigli, Simone Paolo Ponzetto |
Artif. Intell. | 1 |
| 2013 | Collaboratively built semi-structured content and Artificial Intelligence: The story so far
Eduard H. Hovy, Roberto Navigli, Simone Paolo Ponzetto |
Artif. Intell. | 1 |
| 2013 | What Is a Paraphrase?abstractParaphrases are sentences or phrases that convey the same meaning using different wording. Although the logical definition of paraphrases requires strict semantic equivalence, linguistics accepts a broader, approximate, equivalence—thereby allowing far more examples of “quasi-paraphrase.” But approximate equivalence is hard to define. Thus, the phenomenon of paraphrases, as understood in linguistics, is difficult to characterize. In this article, we list a set of 25 operations that generate quasi-paraphrases. We then empirically validate the scope and accuracy of this list by manually analyzing random samples of two publicly available paraphrase corpora. We provide the distribution of naturally occurring quasi-paraphrases in English text. Rahul Bhagat, Eduard H. Hovy |
Comput. Linguistics | 2 |
| 2012 | Evaluating Machine Reading Systems through Comprehension Tests
Anselmo Peñas, Eduard H. Hovy, Pamela Forner, Álvaro Rodrigo, Richard F. E. Sutcliffe, Corina Forascu, Caroline Sporleder |
LREC | 2 |
| 2012 | Structured Event Retrieval over Microblog Archives
Donald Metzler, Congxing Cai, Eduard H. Hovy |
HLT-NAACL | 3 |
| 2011 | Unsupervised Discovery of Domain-Specific Knowledge from Text
Dirk Hovy, Chunliang Zhang, Eduard H. Hovy, Anselmo Peñas |
ACL | 3 |
| 2011 | Insights from Network Structure for Text Mining
Zornitsa Kozareva, Eduard H. Hovy |
ACL | 2 |
| 2011 | A Fast, Accurate, Non-Projective, Semantically-Enriched Parser
Stephen Tratz, Eduard H. Hovy |
EMNLP | 2 |
| 2011 | Causal markers across domains and genres of discourseabstractThis paper is a study of causation as it occurs in different domains and genres of discourse. There have been various initiatives to extract causality from discourse using causal markers. However, to our knowledge, none of these approaches have displayed similar results when applied to other styles of discourse. In this study we evaluate the nature of causal markers - specifically causatives, between corpora in different domains and genres of discourse and measure the overlap of causal markers using two metrics - Term Similarity and Causal Precision. We find that causal markers, specially causatives (causal verbs) are extremely domain dependent, and moderately genre dependent. Rutu Mulkar-Mehta, Andrew S. Gordon, Jerry R. Hobbs, Eduard H. Hovy |
K-CAP | 4 |
| 2011 | Knowledge Engineering Tools for Reasoning with Scientific Observations and Interpretations: a Neural Connectivity Use CaseabstractBACKGROUND: We address the goal of curating observations from published experiments in a generalizable form; reasoning over these observations to generate interpretations and then querying this interpreted knowledge to supply the supporting evidence. We present web-application software as part of the 'BioScholar' project (R01-GM083871) that fully instantiates this process for a well-defined domain: using tract-tracing experiments to study the neural connectivity of the rat brain. RESULTS: The main contribution of this work is to provide the first instantiation of a knowledge representation for experimental observations called 'Knowledge Engineering from Experimental Design' (KEfED) based on experimental variables and their interdependencies. The software has three parts: (a) the KEfED model editor - a design editor for creating KEfED models by drawing a flow diagram of an experimental protocol; (b) the KEfED data interface - a spreadsheet-like tool that permits users to enter experimental data pertaining to a specific model; (c) a 'neural connection matrix' interface that presents neural connectivity as a table of ordinal connection strengths representing the interpretations of tract-tracing data. This tool also allows the user to view experimental evidence pertaining to a specific connection. BioScholar is built in Flex 3.5. It uses Persevere (a noSQL database) as a flexible data store and PowerLoom® (a mature First Order Logic reasoning system) to execute queries using spatial reasoning over the BAMS neuroanatomical ontology. CONCLUSIONS: We first introduce the KEfED approach as a general approach and describe its possible role as a way of introducing structured reasoning into models of argumentation within new models of scientific publication. We then describe the design and implementation of our example application: the BioScholar software. This is presented as a possible biocuration interface and supplementary reasoning toolkit for a larger, more specialized bioinformatics system: the Brain Architecture Management System (BAMS). Thomas A. Russ, Cartic Ramakrishnan, Eduard H. Hovy, Mihail Bota, Gully A. P. C. Burns |
BMC Bioinform. | 3 |
| 2011 | BLANC: Implementing the Rand index for coreference evaluationabstractAbstract This paper addresses the current state of coreference resolution evaluation, in which different measures (notably, MUC, B 3 , CEAF, and ACE-value) are applied in different studies. None of them is fully adequate, and their measures are not commensurate. We enumerate the desiderata for a coreference scoring measure, discuss the strong and weak points of the existing measures, and propose the BiLateral Assessment of Noun-Phrase Coreference, a variation of the Rand index created to suit the coreference task. The BiLateral Assessment of Noun-Phrase Coreference rewards both coreference and non-coreference links by averaging the F-scores of the two types, does not ignore singletons – the main problem with the MUC score – and does not inflate the score in their presence – a problem with the B 3 and CEAF scores. In addition, its fine granularity is consistent over the whole range of scores and affords better discrimination between systems. Marta Recasens, Eduard H. Hovy |
Nat. Lang. Eng. | 2 |
| 2010 | Learning Arguments and Supertypes of Semantic Relations Using Recursive Patterns
Zornitsa Kozareva, Eduard H. Hovy |
ACL | 2 |
| 2010 | Coreference Resolution across Corpora: Languages, Coding Schemes, and Preprocessing Information
Marta Recasens, Eduard H. Hovy |
ACL | 2 |
| 2010 | A Taxonomy, Dataset, and Classifier for Automatic Noun Compound Interpretation
Stephen Tratz, Eduard H. Hovy |
ACL | 2 |
| 2010 | A Semi-Supervised Method to Learn and Construct Taxonomies Using the Web
Zornitsa Kozareva, Eduard H. Hovy |
EMNLP | 2 |
| 2010 | A Typology of Near-Identity Relations for Coreference (NIDENT)
Marta Recasens, Eduard H. Hovy, Maria Antònia Martí |
LREC | 2 |
| 2010 | Not All Seeds Are Equal: Measuring the Quality of Text Mining Seeds
Zornitsa Kozareva, Eduard H. Hovy |
HLT-NAACL | 2 |
| 2010 | Annotation and verification of sense pools in OntoNotes
Liang-Chih Yu, Chung-Hsien Wu 0001, Ru-Yng Chang, Chao-Hong Liu, Eduard H. Hovy |
Inf. Process. Manag. | 5 |
| 2010 | Interlingual annotation of parallel text corpora: a new framework for annotation and evaluationabstractAbstract This paper focuses on an important step in the creation of a system of meaning representation and the development of semantically annotated parallel corpora, for use in applications such as machine translation, question answering, text summarization, and information retrieval. The work described below constitutes the first effort of any kind to annotate multiple translations of foreign-language texts with interlingual content. Three levels of representation are introduced: deep syntactic dependencies (IL0), intermediate semantic representations (IL1), and a normalized representation that unifies conversives, nonliteral language, and paraphrase (IL2). The resulting annotated, multilingually induced, parallel corpora will be useful as an empirical basis for a wide range of research, including the development and evaluation of interlingual NLP systems and paraphrase-extraction systems as well as a host of other research and development efforts in theoretical and applied linguistics, foreign language pedagogy, translation studies, and other related disciplines. Bonnie J. Dorr, Rebecca J. Passonneau, David Farwell, Rebecca Green, Nizar Habash, Stephen Helmreich, Eduard H. Hovy, Lori S. Levin, Keith J. Miller, Teruko Mitamura, Owen Rambow, Advaith Siddharthan |
Nat. Lang. Eng. | 7 |
| 2010 | Introduction to the Special Issue
Eduard H. Hovy, Susie Stephens |
J. Web Semant. | 1 |
| 2009 | Toward Completeness in Concept Extraction and Classification
Eduard H. Hovy, Zornitsa Kozareva, Ellen Riloff |
EMNLP | 1 |
| 2009 | Acquiring paraphrases from text corporaabstractParaphrases are textual expressions that convey the same meaning using different surface forms. Capturing the variability of language, they play an important role in many natural language applications includ ing question answering, machine translation, and multi-document summarization. In linguistics, paraphrases are characterized by approximate conceptual equivalence. Since no automated semantic interpretation systems available today can identify conceptual equivalence, paraphrases are difficult to acquire without human effort. In this paper, we present a method for automatically acquiring paraphrases using a monolingual corpus. We learn paraphrases at both the surface and lexico-syntactic levels and build two paraphrase resources each containing about 2 million phrases. We evaluate these paraphrases extrinsically by using them to learn patterns for Information Extraction (IE). We show that the lexico-syntactic paraphrases performs better than the surface-level paraphrases for IE. We further show that the patterns learned using the lexico-syntactic paraphrases attain comparable performance to the traditional IE approach of learning patterns from domain-specific corpora. Rahul Bhagat, Eduard H. Hovy, Siddharth Patwardhan |
K-CAP | 2 |
| 2009 | Turning the Web into a Database: Extracting Data and Structure
Eduard H. Hovy |
NLDB | 1 |
| 2008 | Semantic Class Learning from the Web with Hyponym Pattern Linkage Graphs
Zornitsa Kozareva, Ellen Riloff, Eduard H. Hovy |
ACL | 3 |
| 2008 | OntoNotes: Corpus Cleanup of Mistaken Agreement Using Word Sense Disambiguation
Liang-Chih Yu, Chung-Hsien Wu 0001, Eduard H. Hovy |
COLING | 3 |
| 2008 | Multi-Criteria-Based Strategy to Stop Active Learning for Data Annotation
Huizhen Wang, Eduard H. Hovy |
COLING | 3 |
| 2008 | Towards Automated Semantic Analysis on Biomedical Research Articles
Donghui Feng 0001, Gully A. P. C. Burns, Eduard H. Hovy |
IJCNLP | 4 |
| 2008 | Learning a Stopping Criterion for Active Learning for Word Sense Disambiguation and Text Classification
Huizhen Wang, Eduard H. Hovy |
IJCNLP | 3 |
| 2008 | A Common Ground for Virtual Humans: Using an Ontology in a Natural Language Oriented Virtual Human Architecture
Arno Hartholt, Thomas A. Russ, David R. Traum, Eduard H. Hovy, Susan Robinson |
LREC | 4 |
| 2008 | The Dynamic Web Presentations with a Generality Model on the News Domain
Hyun Woong Shin, Eduard H. Hovy, Dennis McLeod |
SOFSEM | 2 |
| 2007 | Learning by Reading: A Prototype System, Performance Baseline and Lessons Learned
Ken Barker 0002, Bhalchandra Agashe, Shaw Yi Chaw, James Fan, Noah S. Friedland, Michael R. Glass, Jerry R. Hobbs, Eduard H. Hovy, David J. Israel, Doo Soon Kim, Rutu Mulkar-Mehta, Sourabh Patwardhan, Bruce W. Porter, Dan Tecuci, Peter Z. Yeh |
AAAI | 8 |
| 2007 | Topic Analysis for Psychiatric Document Retrieval
Liang-Chih Yu, Chung-Hsien Wu 0001, Chin-Yew Lin, Eduard H. Hovy, Chia-Ling Lin |
ACL | 4 |
| 2007 | LEDIR: An Unsupervised Algorithm for Learning Directionality of Inference Rules
Rahul Bhagat, Patrick Pantel, Eduard H. Hovy |
EMNLP-CoNLL | 3 |
| 2007 | Extracting Data Records from Unstructured Biomedical Full Text
Donghui Feng 0001, Gully A. P. C. Burns, Eduard H. Hovy |
EMNLP-CoNLL | 3 |
| 2007 | Crystal: Analyzing Predictive Opinions on the Web
Eduard H. Hovy |
EMNLP-CoNLL | 2 |
| 2007 | Active Learning for Word Sense Disambiguation with Methods for Addressing the Class Imbalance Problem
Eduard H. Hovy |
EMNLP-CoNLL | 2 |
| 2007 | Phonetic Models for Generating Spelling Variants
Rahul Bhagat, Eduard H. Hovy |
IJCAI | 2 |
| 2007 | Information acquisition using multiple classificationsabstractGiven a large collection of documents, we often need to extract various aspects of information that may be integrated to form a coherent overall picture. Especially for subjective documents addressing a single topic, traditional summarization techniques are limited in differentiating and clustering similar information. We apply multiple classifications to handle diverse aspects, including subtopic identification, keyword extraction, argument structure analysis, and opinion classification, in order to provide a summarized overview of the collection, complete with distributional information. From this overall summary, system users can effectively obtain more fine-grained information. Our methods for individual modules significantly outperform the baseline and achieve human-level agreement. Namhee Kwon, Eduard H. Hovy |
K-CAP | 2 |
| 2007 | ISP: Learning Inferential Selectional Preferences
Patrick Pantel, Rahul Bhagat, Bonaventura Coppola, Timothy Chklovski, Eduard H. Hovy |
HLT-NAACL | 5 |
| 2006 | Towards Modeling Threaded Discussions using Induced Ontology Knowledge
Donghui Feng 0001, Jihie Kim, Erin Shaw, Eduard H. Hovy |
AAAI | 4 |
| 2006 | Mining and Re-ranking for Answering Biographical Queries on the Web
Donghui Feng 0001, Deepak Ravichandran, Eduard H. Hovy |
AAAI | 3 |
| 2006 | Automatic Identification of Pro and Con Reasons in Online Reviews
Eduard H. Hovy |
ACL | 2 |
| 2006 | Authoring and Generation of Individualized Patient Education Materials
Chrysanne Di Marco, Peter Bray, H. Dominic Covvey, Donald D. Cowan, Vic Di Ciccio, Eduard H. Hovy, Joan Lipa |
AMIA | 6 |
| 2006 | Integrating Semantic Frames from Multiple Sources
Namhee Kwon, Eduard H. Hovy |
CICLing | 2 |
| 2006 | Re-evaluating Machine Translation Results with Paraphrase Support
Chin-Yew Lin, Eduard H. Hovy |
EMNLP | 3 |
| 2006 | Toward Large-Scale Shallow Semantics for Higher-Quality NLP
Eduard H. Hovy |
ESWC | 1 |
| 2006 | An intelligent discussion-bot for answering student queries in threaded discussionsabstractThis paper describes a discussion-bot that provides answers to students' discussion board questions in an unobtrusive and human-like way. Using information retrieval and natural language processing techniques, the discussion-bot identifies the questioner's interest, mines suitable answers from an annotated corpus of 1236 archived threaded discussions and 279 course documents and chooses an appropriate response. A novel modeling approach was designed for the analysis of archived threaded discussions to facilitate answer extraction. We compare a self-out and an all-in evaluation of the mined answers. The results show that the discussion-bot can begin to meet students' learning requests. We discuss directions that might be taken to increase the effectiveness of the question matching and answer extraction algorithms. The research takes place in the context of an undergraduate computer science course. Donghui Feng 0001, Erin Shaw, Jihie Kim, Eduard H. Hovy |
IUI | 4 |
| 2006 | Taking advantage of the situation: non-linguistic context for natural language interfaces to interactive virtual environmentsabstractWe introduce a framework for learning situated Natural Language Interfaces (NLIs) to interactive virtual environments. The framework exploits the non-linguistic context, or situation, explicitly modeled in such interactive applications. This situation model is integrated with a model of word meaning in a principled manner using a noisy channel approach to language understanding. Preliminary experimentation in an independently designed interactive application, i.e. the Mission Rehearsal Exercise (MRE), shows that this situated NLI outperforms a state of the art NLI on both whole frame accuracy and F-Score metrics. Further, use of the situation model in the situated NLI is shown to increase robustness to the noise introduced by the use of automatic speech recognition. Michael Fleischman, Eduard H. Hovy |
IUI | 2 |
| 2006 | Automated Summarization Evaluation with Basic Elements
Eduard H. Hovy, Chin-Yew Lin, Jun-ichi Fukumoto |
LREC | 1 |
| 2006 | Parallel Syntactic Annotation of Multiple Languages
Owen Rambow, Bonnie J. Dorr, David Farwell, Rebecca Green, Nizar Habash, Stephen Helmreich, Eduard H. Hovy, Lori S. Levin, Keith J. Miller, Teruko Mitamura, Flo Reeder, Advaith Siddharthan |
LREC | 7 |
| 2006 | Summarizing Answers for Complicated Questions
Chin-Yew Lin, Eduard H. Hovy |
LREC | 3 |
| 2006 | Learning to Detect Conversation Focus of Threaded Discussions
Donghui Feng 0001, Erin Shaw, Jihie Kim, Eduard H. Hovy |
HLT-NAACL | 4 |
| 2006 | OntoNotes: The 90% Solution
Eduard H. Hovy, Mitchell P. Marcus, Martha Palmer, Lance A. Ramshaw, Ralph M. Weischedel |
HLT-NAACL | 1 |
| 2006 | Identifying and Analyzing Judgment Opinions
Eduard H. Hovy |
HLT-NAACL | 2 |
| 2006 | ParaEval: Using Paraphrases to Evaluate Summaries Automatically
Chin-Yew Lin, Dragos Stefan Munteanu, Eduard H. Hovy |
HLT-NAACL | 4 |
| 2005 | Randomized Algorithms and NLP: Using Locality Sensitive Hash Functions for High Speed Noun ClusteringabstractIn this paper, we explore the power of randomized algorithm to address the challenge of working with very large amounts of data. We apply these algorithms to generate noun similarity lists from 70 million pages. We reduce the running time from quadratic to practically linear in the number of elements to be computed. Deepak Ravichandran, Patrick Pantel, Eduard H. Hovy |
ACL | 3 |
| 2005 | Digesting Virtual "Geek" Culture: The Summarization of Technical Internet Relay ChatsabstractThis paper describes a summarization system for technical chats and emails on the Linux kernel. To reflect the complexity and sophistication of the discussions, they are clustered according to subtopic structure on the sub-message level, and immediate responding pairs are identified through machine learning methods. A resulting summary consists of one or more mini-summaries, each on a subtopic from the discussion. Eduard H. Hovy |
ACL | 2 |
| 2005 | An Information Theoretic Model for Database Alignment
Patrick Pantel, Andrew Philpot, Eduard H. Hovy |
SSDBM | 3 |
| 2005 | A Framework for Effective Annotation of Information from Closed Captions Using Ontologies
Latifur Khan, Dennis McLeod, Eduard H. Hovy |
J. Intell. Inf. Syst. | 3 |
| 2004 | Determining the Sentiment of Opinions
Eduard H. Hovy |
COLING | 2 |
| 2004 | FrameNet-based Semantic Parsing using Maximum Entropy Models
Namhee Kwon, Michael Fleischman, Eduard H. Hovy |
COLING | 3 |
| 2004 | Towards Terascale Semantic Acquisition
Patrick Pantel, Deepak Ravichandran, Eduard H. Hovy |
COLING | 3 |
| 2004 | Multi-Document Biography Summarization
Miruna Ticrea, Eduard H. Hovy |
EMNLP | 3 |
| 2004 | Retrieval effectiveness of an ontology-based model for information selection
Latifur Khan, Dennis McLeod, Eduard H. Hovy |
VLDB J. | 3 |
| 2003 | Offline Strategies for Online Question Answering: Answering Questions Before They Are AskedabstractRecent work in Question Answering has focused on web-based systems that extract answers using simple lexico-syntactic patterns. We present an alternative strategy in which patterns are used to extract highly precise relational information offline, creating a data repository that is used to efficiently answer questions. We evaluate our strategy on a challenging subset of questions, i.e. "Who is ..." questions, against a state of the art web-based Question Answering system. Results indicate that the extracted relations answer 25% more questions correctly and do so three orders of magnitude faster than the state of the art system. Michael Fleischman, Eduard H. Hovy, Abdessamad Echihabi |
ACL | 2 |
| 2003 | Maximum Entropy Models for FrameNet Classification
Michael Fleischman, Namhee Kwon, Eduard H. Hovy |
EMNLP | 3 |
| 2003 | Recommendations without user preferences: a natural language processing approachabstractWe examine the problems with automated recommendation systems when information about user preferences is limited. We equate the problem to one of content similarity measurement and apply techniques from Natural Language Processing to the domain of movie recommendation. We describe two algorithms, a naïve word-space approach and a more sophisticated approach using topic signatures, and evaluate their performance compared to baseline, gold standard, and commercial systems. Michael Fleischman, Eduard H. Hovy |
IUI | 2 |
| 2003 | BTL: a hybrid model for English-Vietnamese machine translationabstractMachine Translation (MT) is the most interesting and difficult task which has been posed since the beginning of computer history. The highest difficulty which computers had to face with, is the built-in ambiguity of Natural Languages. Formerly, a lot of human-devised rules have been used to disambiguate those ambiguities. Building such a complete rule-set is time-consuming and labor-intensive task whilst it doesn’t cover all the cases. Besides, when the scale of system increases, it is very difficult to control that rule-set. In this paper, we present a new model of learning-based MT (entitled BTL: Bitext-Transfer Learning) that learns from bilingual corpus to extract disambiguating rules. This model has been experimented in English-to-Vietnamese MT system (EVT) and it gave encouraging results. Dinh Dien, Kiem Hoang, Eduard H. Hovy |
MTSummit | 3 |
| 2003 | FEMTI: creating and using a framework for MT evaluationabstractThis paper presents FEMTI, a web-based Framework for the Evaluation of Machine Translation in ISLE. FEMTI offers structured descriptions of potential user needs, linked to an overview of technical characteristics of MT systems. The description of possible systems is mainly articulated around the quality characteristics for software product set out in ISO/IEC standard 9126. Following the philosophy set out there and in the related 14598 series of standards, each quality characteristic bottoms out in metrics which may be applied to a particular instance of a system in order to judge how satisfactory the system is with respect to that characteristic. An evaluator can use the description of user needs to help identify the specific needs of his evaluation and the relations between them. He can then follow the pointers to system description to determine what metrics should be applied and how. In the current state of the framework, emphasis is on being exhaustive, including as much as possible of the information available in the literature on machine translation evaluation. Future work will aim at being more analytic, looking at characteristics and metrics to see how they relate to one another, validating metrics and investigating the correlation between particular metrics and human judgement. Margaret King, Andrei Popescu-Belis, Eduard H. Hovy |
MTSummit | 3 |
| 2003 | A Maximum Entropy Approach to FrameNet Tagging
Michael Fleischman, Eduard H. Hovy |
HLT-NAACL | 2 |
| 2003 | Automatic Evaluation of Summaries Using N-gram Co-occurrence Statistics
Chin-Yew Lin, Eduard H. Hovy |
HLT-NAACL | 2 |
| 2003 | A Web-Trained Extraction Summarization System
Eduard H. Hovy |
HLT-NAACL | 2 |
| 2003 | Cross-lingual C*ST*RD: English access to Hindi informationabstractWe present C*ST*RD, a cross-language information delivery system that supports cross-language information retrieval, information space visualization and navigation, machine translation, and text summarization of single documents and clusters of documents. C*ST*RD was assembled and trained within 1 month, in the context of DARPA's Surprise Language Exercise, that selected as source a heretofore unstudied language, Hindi. Given the brief time, we could not create deep Hindi capabilities for all the modules, but instead experimented with combining shallow Hindi capabilities, or even English-only modules, into one integrated system. Various possible configurations, with different tradeoffs in processing speed and ease of use, enable the rapid deployment of C*ST*RD to new languages under various conditions. Anton Leuski, Chin-Yew Lin, Ulrich Germann, Franz Josef Och, Eduard H. Hovy |
ACM Trans. Asian Lang. Inf. Process. | 6 |
| 2002 | From Single to Multi-document SummarizationabstractNeATS is a multi-document summarization system that attempts to extract relevant or interesting portions from a set of documents about some topic and present them in coherent order. NeATS is among the best performers in the large scale summarization evaluation DUC 2001. Chin-Yew Lin, Eduard H. Hovy |
ACL | 2 |
| 2002 | Learning surface text patterns for a Question Answering SystemabstractIn this paper we explore the power of surface text patterns for open-domain question answering systems. In order to obtain an optimal set of patterns, we have developed a method for learning such patterns automatically. A tagged corpus is built from the Internet in a bootstrapping process by providing a few hand-crafted examples of each question type to Altavista. Patterns are then automatically extracted from the returned documents and standardized. We calculate the precision of each pattern, and the average precision for each question type. These patterns are then applied to find answers to new questions. Using the TREC-10 question set, we report results for two cases: answers determined from the TREC-10 corpus and from the web. Deepak Ravichandran, Eduard H. Hovy |
ACL | 2 |
| 2002 | Fine Grained Classification of Named Entities
Michael Fleischman, Eduard H. Hovy |
COLING | 2 |
| 2002 | Using Knowledge to Facilitate Factoid Answer Pinpointing
Eduard H. Hovy, Ulf Hermjakob, Chin-Yew Lin, Deepak Ravichandran |
COLING | 1 |
| 2002 | Towards Emotional Variation in Speech-Based Natural Language Processing
Michael Fleischman, Eduard H. Hovy |
INLG | 2 |
| 2002 | Computer-Aided Specification of Quality Models for Machine Translation Evaluation
Eduard H. Hovy, Margaret King, Andrei Popescu-Belis |
LREC | 1 |
| 2002 | Learning, Collecting, and Using Ontological Knowledge for NLP
Eduard H. Hovy |
PRICAI | 1 |
| 2002 | Introduction to the Special Issue on Summarizationabstractgeneration based on rhetorical structure extraction. In Proceedings of the International Conference on Computational Linguistics, Kyoto, Japan, pages 344–348. Otterbacher, Jahna, Dragomir R. Radev, and Airong Luo. 2002. Revisions that improve cohesion in multi-document summaries: A preliminary study. In ACL Workshop on Text Summarization, Philadelphia. Papineni, K., S. Roukos, T. Ward, and W-J. Zhu. 2001. BLEU: A method for automatic evaluation of machine translation. Research Report RC22176, IBM. Radev, Dragomir, Simone Teufel, Horacio Saggion, Wai Lam, John Blitzer, Arda Celebi, Hong Qi, Elliott Drabek, and Danyu Liu. 2002. Evaluation of text summarization in a cross-lingual information retrieval framework. Technical Report, Center for Language and Speech Processing, Johns Hopkins University, Baltimore, June. Radev, Dragomir R., Hongyan Jing, and Malgorzata Budzikowska. 2000. Centroid-based summarization of multiple documents: Sentence extraction, utility-based evaluation, and user studies. In ANLP/NAACL Workshop on Summarization, Seattle, April. Radev, Dragomir R. and Kathleen R. McKeown. 1998. Generating natural language summaries from multiple on-line sources. Computational Linguistics, 24(3):469–500. Rau, Lisa and Paul Jacobs. 1991. Creating segmented databases from free text for text retrieval. In Proceedings of the 14th Annual International ACM-SIGIR Conference on Research and Development in Information Retrieval, New York, pages 337–346. Saggion, Horacio and Guy Lapalme. 2002. Generating indicative-informative summaries with SumUM. Computational Linguistics, 28(4), 497–526. Salton, G., A. Singhal, M. Mitra, and C. Buckley. 1997. Automatic text structuring and summarization. Information Processing & Management, 33(2):193–207. Silber, H. Gregory and Kathleen McCoy. 2002. Efficiently computed lexical chains as an intermediate representation for automatic text summarization. Computational Linguistics, 28(4), 487–496. Sparck Jones, Karen. 1999. Automatic summarizing: Factors and directions. In I. Mani and M. T. Maybury, editors, Advances in Automatic Text Summarization. MIT Press, Cambridge, pages 1–13. Strzalkowski, Tomek, Gees Stein, J. Wang, and Bowden Wise. 1999. A robust practical text summarizer. In I. Mani and M. T. Maybury, editors, Advances in Automatic Text Summarization. MIT Press, Cambridge, pages 137–154. Teufel, Simone and Marc Moens. 2002. Summarizing scientific articles: Experiments with relevance and rhetorical status. Computational Linguistics, 28(4), 409–445. White, Michael and Claire Cardie. 2002. Selecting sentences for multidocument summaries using randomized local search. In Proceedings of the Workshop on Automatic Summarization (including DUC 2002), Philadelphia, July. Association for Computational Linguistics, New Brunswick, NJ, pages 9–18. Witbrock, Michael and Vibhu Mittal. 1999. Ultra-summarization: A statistical approach to generating highly condensed non-extractive summaries. In Proceedings of the 22nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Berkeley, pages 315–316. Zechner, Klaus. 2002. Automatic summarization of open-domain multiparty dialogues in diverse genres. Computational Linguistics, 28(4), 447–485. Dragomir R. Radev, Eduard H. Hovy, Kathy McKeown |
Comput. Linguistics | 2 |
| 2002 | Principles of Context-Based Machine Translation Evaluation
Eduard H. Hovy, Margaret King, Andrei Popescu-Belis |
Mach. Transl. | 1 |
| 2000 | The Automated Acquisition of Topic Signatures for Text Summarization
Chin-Yew Lin, Eduard H. Hovy |
COLING | 2 |
| 1999 | MT evaluationabstractThis panel deals with the general topic of evaluation of machine translation systems. The first contribution sets out some recent work on creating standards for the design of evaluations. The second, by Eduard Hovy. takes up the particular issue of how metrics can be differentiated and systematized. Benjamin K. T’sou suggests that whilst men may evaluate machines, machines may also evaluate men. John S. White focuses on the question of the role of the user in evaluation design, and Yusoff Zaharin points out that circumstances and settings may have a major influence on evaluation design. Margaret King, Eduard H. Hovy, Benjamin Ka-Yin T'sou, Yusoff Zaharin |
MTSummit | 2 |
| 1998 | Combining and standardizing large- scale, practical ontologies for machine tranlation and other uses
Eduard H. Hovy |
LREC | 1 |
| 1996 | The HealthDoc Sentence PlannerabstractThis paper describes the Sentence Planner (sP) in the HealthDoc project, which is concerned with the production of customized patienteducation material from a source encoded in terms of plans.The task of the sP is to transform selected, not necessarily consecutive, plans (which may vary in detail, from text plans specifying only content and discourse organization to fine-grained but incohesive, sentence plans) into completely specified specifications for the surface generator.The paper identifies the sentence planning tasks, which are highly interdependent and partially parallel, and argues, in accordance with [Nirenburg et al., 1989], that' a blackboard architecture with several independent modules is most suitable to deal with them.The architecture is presented, and the interaction of the sentence planning modules within this architecture is shown.The first implementation of the sP is discussed; examples illustrate the planning process in action. Leo Wanner, Eduard H. Hovy |
INLG (1) | 2 |
| 1995 | Filling Knowledge Gaps in a Broad-Coverage Machine Translation System
Kevin Knight, Ishwar Chander, Matthew Haines, Vasileios Hatzivassiloglou, Eduard H. Hovy, Masayo Iida, Steve K. Luk, Richard Whitney, Kenji Yamada |
IJCAI | 5 |
| 1994 | Toward a Multidimensional Framework to Guide the Automated Generation of Text Types
Julia Lavid, Eduard H. Hovy |
INLG | 2 |
| 1993 | Structure and Rules in Automated Multimedia Presentation Planning
Yigal Arens, Eduard H. Hovy, Susanne van Mulken |
IJCAI | 2 |
| 1993 | Automated Discourse Generation Using Discourse Structure Relations
Eduard H. Hovy |
Artif. Intell. | 1 |
| 1993 | Good applications for crummy machine translation
Kenneth Church 0001, Eduard H. Hovy |
Mach. Transl. | 2 |
| 1991 | Automatic Generation of Formatted Text
Eduard H. Hovy, Yigal Arens |
AAAI | 1 |
| 1990 | Parsimonious and Profligate Approaches to the Question of Discourse Structure Relations
Eduard H. Hovy |
INLG | 1 |
| 1990 | Pragmatics and Natural Language Generation
Eduard H. Hovy |
Artif. Intell. | 1 |
| 1988 | Planning Coherent Multisentential TextabstractThough most text generators are capable of simply stringing together more than one sentence, they cannot determine which order will ensure a coherent paragraph. A paragraph is coherent when the information in successive sentences follows some pattern of inference or of knowledge with which the hearer is familiar. To signal such inferences, speakers usually use relations that link successive sentences in fixed ways. A set of 20 relations that span most of what people usually say in English is proposed in the Rhetorical Structure Theory of Mann and Thompson. This paper describes the formalization of these relations and their use in a prototype text planner that structures input elements into coherent paragraphs. Eduard H. Hovy |
ACL | 1 |
| 1988 | Two Types of Planning in Language GenerationabstractAs our understanding of natural language generation has increased, a number of tasks have been separated from realization and put together under the heading "text planning". So far, however, no-one has enumerated the kinds of tasks a text planner should be able to do. This paper describes the principal lesson learned in combining a number of planning tasks in a planner-realizer: planning and realization should be interleaved, in a limited-commitment planning paradigm, to perform two types of planning: prescriptive and restrictive. Limited-commitment planning consists of both prescriptive (hierarchical expansion) planning and of restrictive planning (selecting from options with reference to the status of active goals). At present, existing text planners use prescriptive plans exclusively. However, a large class of planner tasks, especially those concerned with the pragmatic (non-literal) content of text such as style and slant, is most easily performed under restrictive planning. The kinds of tasks suited to each planning style are listed, and a program that uses both styles is described. Eduard H. Hovy |
ACL | 1 |
| 1987 | Interpretation in Generation
Eduard H. Hovy |
AAAI | 1 |
| 1985 | Integrating Text Planning and Production in Generation
Eduard H. Hovy |
IJCAI | 1 |