EDBT 2026 Demo / reviewers in the wild / expert
Bonnie J. Dorr
dblp:d/BonnieJDorr · also Bonnie Dorr
· DBLP profile ↗
86ranked-venue papers
31as first author
9since 2021 · last 2026
0000-0003-4356-5813ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 75 · 30 first-author · 8 since 2021Databases, data management, data science and information retrieval · 16 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multilingual Target-Stance Extraction
Ethan Mines, Bonnie J. Dorr |
LREC | 2 |
| 2025 | Making Task-Oriented Dialogue Datasets More Natural by Synthetically Generating Indirect User RequestsabstractIndirect User Requests (IURs), such as “It’s cold in here” instead of “Could you please increase the temperature?” are common in human-human task-oriented dialogue and require world knowledge and pragmatic reasoning from the listener. While large language models (LLMs) can handle these requests effectively, smaller models deployed on virtual assistants often struggle due to resource constraints. Moreover, existing task-oriented dialogue benchmarks lack sufficient examples of complex discourse phenomena such as indirectness. To address this, we propose a set of linguistic criteria along with an LLM-based pipeline for generating realistic IURs to test natural language understanding (NLU) and dialogue state tracking (DST) models before deployment in a new domain. We also release IndirectRequests, a dataset of IURs based on the Schema-Guided Dialogue (SGD) corpus, as a comparative testbed for evaluating the performance of smaller models in handling indirect requests. Amogh Mannekote, Jinseok Nam, Kristy Elizabeth Boyer, Bonnie J. Dorr |
COLING | 5 |
| 2025 | Exploiting Explainability to Design Adversarial Attacks and Evaluate Attack Resilience in Hate-Speech Detection ModelsabstractThe advent of social media has given rise to numerous ethical challenges, with hate speech among the most significant concerns. Researchers are attempting to tackle this problem by using hate-speech detection and employing language models to automatically moderate content and promote civil discourse. Unfortunately, recent studies have revealed that hate-speech detection systems can be misled by adversarial attacks, raising concerns about their resilience. While previous research has separately addressed the robustness of these models under adversarial attacks and their explainability, there has been no comprehensive study exploring their intersection. The novelty of our work lies in combining these two critical aspects, leveraging explainability to identify potential vulnerabilities and enabling the design of targeted adversarial attacks. This paper quantifies the interplay between explainability and adversarial robustness in hate-speech detection models. We define novel metrics based on explainability-driven adversarial attacks to evaluate this relationship, providing a clear assessment of model vulnerabilities and guiding the development of more resilient systems. Pranath Reddy Kumbam, Sohaib Uddin Syed, Prashanth Thamminedi, Suhas Harish, Ian Perera, Bonnie J. Dorr |
ICWSM | 6 |
| 2025 | DETQUS: Decomposition-Enhanced Transformers for QUery-focused SummarizationabstractYasir Khan, Xinlei Wu, Sangpil Youm, Justin Ho, Aryaan Mehboob Shaikh, Jairo Garciga, Rohan Sharma, Bonnie J Dorr. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yasir Khan, Xinlei Wu, Sangpil Youm, Justin Ho, Aryaan Shaikh, Jairo Garciga, Rohan Sharma, Bonnie J. Dorr |
NAACL (Long Papers) | 8 |
| 2024 | The Effect of Data Partitioning Strategy on Model Generalizability: A Case Study of Morphological SegmentationabstractRecent work to enhance data partitioning strategies for more realistic model evaluation face challenges in providing a clear optimal choice.This study addresses these challenges, focusing on morphological segmentation and synthesizing limitations related to language diversity, adoption of multiple datasets and splits, and detailed model comparisons.Our study leverages data from 19 languages, including ten indigenous or endangered languages across 10 language families with diverse morphological systems (polysynthetic, fusional, and agglutinative) and different degrees of data availability.We conduct large-scale experimentation with varying sized combinations of training and evaluation sets as well as new test data.Our results show that, when faced with new test data: (1) models trained from random splits are able to achieve higher numerical scores; (2) model rankings derived from random splits tend to generalize more consistently. Zoey Liu, Bonnie J. Dorr |
NAACL-HLT | 2 |
| 2024 | DAHRS: Divergence-Aware Hallucination-Remediated SRL Projection
Sangpil Youm, Brodie Mather, Chathuri Jayaweera, Juliana Prada, Bonnie J. Dorr |
NLDB (1) | 5 |
| 2023 | LonXplain: Lonesomeness as a Consequence of Mental Disturbance in Reddit Posts
Muskan Garg, Chandni Saxena, Debabrata Samanta, Bonnie J. Dorr |
NLDB | 4 |
| 2022 | BeSt: The Belief and Sentiment CorpusabstractWe present the BeSt corpus, which records cognitive state: who believes what (i.e., factuality), and who has what sentiment towards what. This corpus is inspired by similar source-and-target corpora, specifically MPQA and FactBank. The corpus comprises two genres, newswire and discussion forums, in three languages, Chinese (Mandarin), English, and Spanish. The corpus is distributed through the LDC. Jennifer Tracey, Owen Rambow, Claire Cardie, Adam Dalton 0001, Hoa Trang Dang, Mona T. Diab, Bonnie J. Dorr, Louise Guthrie, Magdalena Markowska, Smaranda Muresan, Vinodkumar Prabhakaran, Samira Shaikh, Tomek Strzalkowski |
LREC | 7 |
| 2021 | Strengthening Low-resource Neural Machine Translation through Joint Learning: The Case of Farsi-Spanish
Benyamin Ahmadnia, Raúl Aranovich, Bonnie J. Dorr |
ICAART (1) | 3 |
| 2020 | Detecting Asks in Social Engineering Attacks: Impact of Linguistic and Structural KnowledgeabstractSocial engineers attempt to manipulate users into undertaking actions such as downloading malware by clicking links or providing access to money or sensitive information. Natural language processing, computational sociolinguistics, and media-specific structural clues provide a means for detecting both the ask (e.g., buy gift card) and the risk/reward implied by the ask, which we call framing (e.g., lose your job, get a raise). We apply linguistic resources such as Lexical Conceptual Structure to tackle ask detection and also leverage structural clues such as links and their proximity to identified asks to improve confidence in our results. Our experiments indicate that the performance of ask detection, framing detection, and identification of the top ask is improved by linguistically motivated classes coupled with structural clues such as links. Our approach is implemented in a system that informs users about social engineering risk situations. Bonnie J. Dorr, Archna Bhatia, Adam Dalton 0001, Brodie Mather, Bryanna Hebenstreit, Sashank Santhanam, Samira Shaikh, Alan Zemel, Tomek Strzalkowski |
AAAI | 1 |
| 2018 | Cyberattack Prediction Through Public Text Analysis and Mini-TheoriesabstractThis paper describes a new approach to detection and tracking of potential cyberattacks from analyzing large quantities of cyber-related webpage text, using ontological knowledge about such attacks combined with composable causal models represented in Probabilistic Soft Logic. The stages of a cyberattack kill chain are viewed as a sequence of both observed and unobserved events (e.g., reconnaissance, weaponize, exploit, install) and explicit mentions of, or related to, such events are examined as potential signals for a future attack. Using a suite of natural language processing techniques, sentences from input news texts are automatically classified according to the described cyberattack event, then enriched with named entity recognition for the rapid detection of key elements that might be associated with potential cyberattacks. We present our work as a framework for rapid and flexible predictive analysis of the ever-increasing amount of cyber-related text data, with initial experiments indicating that event detection using parsing and named entity recognition combined with statistical relational learning show promise in time-series prediction from news text. Ian Perera, Jena D. Hwang, Kevin Bayas, Bonnie J. Dorr, Yorick Wilks |
IEEE BigData | 4 |
| 2017 | Improving cyber-attack predictions through information foragingabstractThis paper describes how information foraging is useful in the implementation of new algorithms to anticipate cyber attacks. The exploration of publicly available data has been used to predict events in the socio-political domain, but the adversarial and covert behavior of actors in cyber security creates additional challenges. This paper describes a framework for Information Foraging for Algorithm Discovery (IFAD) that addresses standard data-science issues of volume and variety, by balancing human intuition with automation, and thus taking initial steps toward supporting the increasing need for rapid analysis of, and tool development for, big data. Our results demonstrate that cognitive augmentation, and information foraging in particular, is useful in the development of tools to anticipate cyber attacks using publicly available data. Adam Dalton 0001, Bonnie J. Dorr, Leon Liang, Kristy Hollingshead |
IEEE BigData | 2 |
| 2016 | Evaluation-driven research in data science: Leveraging cross-field methodologiesabstractWhile prior evaluation methodologies for data-science research have focused on efficient and effective teamwork on independent data science problems within given fields [1], this paper argues that an enriched notion of evaluation-driven research (EDR) supports methodologies and effective solutions to data-science problems across multiple fields. We adopt the view that progress in data-science research is enriched through the examination of a range of problems in many different areas (traffic, healthcare, finance, sports, etc.) and through the development of methodologies and evaluation paradigms that span diverse disciplines, domains, problems, and tasks. A number of questions arise when one considers the multiplicity of data science fields and the potential for cross-disciplinary “sharing” of methodologies, for example: the feasibility of generalizing problems, tasks, and metrics across domains; ground-truth considerations for different types of problems; issues related to data uncertainty in different fields; and the feasibility of enabling cross-field cooperation to encourage diversity of solutions. We posit that addressing the problems inherent in such questions provides a foundation for EDR across diverse fields. We ground our conclusions and insights in a brief preliminary study developed within the Information Access Division of the National Institute of Standards and Technology as a part of a new Data Science Research Program (DSRP). The DSRP focuses on this cross-disciplinary notion of EDR and includes a new Data Science Evaluation series to facilitate research collaboration, to leverage shared technology and infrastructure, and to further build and strengthen the data-science community. Bonnie J. Dorr, Peter C. Fontana, Craig S. Greenberg, Marion Le Bras, Mark A. Przybocki |
IEEE BigData | 1 |
| 2015 | Speech Adaptation in Extended Ambient Intelligence EnvironmentsabstractThis Blue Sky presentation focuses on a major shift toward a notion of “ambient intelligence” that transcends general applications targeted at the general population. The focus is on highly personalized agents that accommodate individual differences and changes over time. This notion of Extended Ambient Intelligence (EAI) concerns adaptation to a person’s preferences and experiences, as well as changing capabilities, most notably in an environment where conversational engagement is central. An important step in moving this research forward is the accommodation of different degrees of cognitive capability (including speech processing) that may vary over time for a given user—whether through improvement or through deterioration. We suggest that the application of divergence detection to speech patterns may enable adaptation to a speaker’s increasing or decreasing level of speech impairment over time. Taking an adaptive approach toward technology development in this arena may be a first step toward empowering those with special needs so that they may live with a high quality of life. It also represents an important step toward a notion of ambient intelligence that is personalized beyond what can be achieved by mass-produced, one-size-fits-all software currently in use on mobile devices. Bonnie J. Dorr, Lucian Galescu, Ian Perera, Kristy Hollingshead, David J. Atkinson 0001, Micah Clark, William J. Clancey, Yorick Wilks, Eric Fosler-Lussier |
AAAI | 1 |
| 2015 | The NIST data science evaluation series: Part of the NIST information access division data science initiativeabstractThe Information Access Division (IAD) of the National Institute of Standards and Technology (NIST) launched a new Data Science Initiative in the fall of 2015. This initiative focuses on evaluation-driven research and will establish a new Data Science Evaluation series to facilitate research collaboration, to leverage shared technology and infrastructure, and to further build and strengthen the data science community. The evaluation series will consist of a pre-pilot to be launched in the fall of 2015, a pilot evaluation to be launched in 2016, and a full-scale multiple-track evaluation in 2017. In addition to these evaluations, this new initiative aims to address several infrastructure challenges and to provide standards to encourage easier group collaboration. Bonnie J. Dorr, Craig S. Greenberg, Peter C. Fontana, Mark A. Przybocki, Marion Le Bras, Cathryn A. Ploehn, Oleg Aulov, Wo Chang |
IEEE BigData | 1 |
| 2015 | The NIST data science initiativeabstractWe examine foundational issues in data science including current challenges, basic research questions, and expected advances, as the basis for a new Data Science Initiative and evaluation series, introduced by the National Institute of Standards and Technology (NIST) in the fall of 2015. The evaluations will facilitate research efforts, collaboration, leverage shared infrastructure, and effectively address cross-cutting challenges faced by diverse data science communities. The evaluations will have multiple research tracks championed by members of the data science community, and will enable rigorous comparison of approaches through common tasks, datasets, metrics, and shared research challenges. The tracks will measure several different data science technologies in a wide range of fields, starting with a pre-pilot. In addition to developing data science evaluation methods and metrics, it will address computing infrastructure, standards for an interoperability framework, and domain-specific examples. Bonnie J. Dorr, Craig S. Greenberg, Peter C. Fontana, Mark A. Przybocki, Marion Le Bras, Cathryn A. Ploehn, Oleg Aulov, Martial Michel, E. Jim Golden, Wo Chang |
DSAA | 1 |
| 2013 | Computing Lexical ContrastabstractKnowing the degree of semantic contrast between words has widespread application in natural language processing, including machine translation, information retrieval, and dialogue systems. Manually created lexicons focus on opposites, such as hot and cold. Opposites are of many kinds such as antipodals, complementaries, and gradable. Existing lexicons often do not classify opposites into the different kinds, however. They also do not explicitly list word pairs that are not opposites but yet have some degree of contrast in meaning, such as warm and cold or tropical and freezing. We propose an automatic method to identify contrasting word pairs that is based on the hypothesis that if a pair of words, A and B, are contrasting, then there is a pair of opposites, C and D, such that A and C are strongly related and B and D are strongly related. (For example, there exists the pair of opposites hot and cold such that tropical is related to hot, and freezing is related to cold.) We will call this the contrast hypothesis. We begin with a large crowdsourcing experiment to determine the amount of human agreement on the concept of oppositeness and its different kinds. In the process, we flesh out key features of different kinds of opposites. We then present an automatic and empirical measure of lexical contrast that relies on the contrast hypothesis, corpus statistics, and the structure of a Roget-like thesaurus. We show how, using four different data sets, we evaluated our approach on two different tasks, solving “most contrasting word” questions and distinguishing synonyms from opposites. The results are analyzed across four parts of speech and across five different kinds of opposites. We show that the proposed measure of lexical contrast obtains high precision and large coverage, outperforming existing methods. Saif M. Mohammad, Bonnie J. Dorr, Graeme Hirst, Peter D. Turney |
Comput. Linguistics | 2 |
| 2013 | Generating Extractive Summaries of Scientific ParadigmsabstractResearchers and scientists increasingly find themselves in the position of having to quickly understand large amounts of technical material. Our goal is to effectively serve this need by using bibliometric text mining and summarization techniques to generate summaries of scientific literature. We show how we can use citations to produce automatically generated, readily consumable, technical extractive summaries. We first propose C-LexRank, a model for summarizing single scientific articles based on citations, which employs community detection and extracts salient information-rich sentences. Next, we further extend our experiments to summarize a set of papers, which cover the same scientific topic. We generate extractive summaries of a set of Question Answering (QA) and Dependency Parsing (DP) papers, their abstracts, and their citation sentences and show that citations have unique information amenable to creating a summary. Vahed Qazvinian, Dragomir R. Radev, Saif M. Mohammad, Bonnie J. Dorr, David M. Zajic, Michael Whidby, Taesun Moon |
J. Artif. Intell. Res. | 4 |
| 2013 | Generating targeted paraphrases for improved translationabstractToday's Statistical Machine Translation (SMT) systems require high-quality human translations for parameter tuning, in addition to large bitexts for learning the translation units. This parameter tuning usually involves generating translations at different points in the parameter space and obtaining feedback against human-authored reference translations as to how good the translations. This feedback then dictates what point in the parameter space should be explored next. To measure this feedback, it is generally considered wise to have multiple (usually 4) reference translations to avoid unfair penalization of translation hypotheses which could easily happen given the large number of ways in which a sentence can be translated from one language to another. However, this reliance on multiple reference translations creates a problem since they are labor intensive and expensive to obtain. Therefore, most current MT datasets only contain a single reference. This leads to the problem of reference sparsity. In our previously published research, we had proposed the first paraphrase-based solution to this problem and evaluated its effect on Chinese-English translation. In this article, we first present extended results for that solution on additional source languages. More importantly, we present a novel way to generate “targeted” paraphrases that yields substantially larger gains (up to 2.7 BLEU points) in translation quality when compared to our previous solution (up to 1.6 BLEU points). In addition, we further validate these improvements by supplementing with human preference judgments obtained via Amazon Mechanical Turk. Nitin Madnani, Bonnie J. Dorr |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2012 | Modality and Negation in SIMT Use of Modality and Negation in Semantically-Informed Syntactic MTabstractThis article describes the resource- and system-building efforts of an 8-week Johns Hopkins University Human Language Technology Center of Excellence Summer Camp for Applied Language Exploration (SCALE-2009) on Semantically Informed Machine Translation (SIMT). We describe a new modality/negation (MN) annotation scheme, the creation of a (publicly available) MN lexicon, and two automated MN taggers that we built using the annotation scheme and lexicon. Our annotation scheme isolates three components of modality and negation: a trigger (a word that conveys modality or negation), a target (an action associated with modality or negation), and a holder (an experiencer of modality). We describe how our MN lexicon was semi-automatically produced and we demonstrate that a structure-based MN tagger results in precision around 86% (depending on genre) for tagging of a standard LDC data set. We apply our MN annotation scheme to statistical machine translation using a syntactic framework that supports the inclusion of semantic annotations. Syntactic tags enriched with semantic annotations are assigned to parse trees in the target-language training texts through a process of tree grafting. Although the focus of our work is modality and negation, the tree grafting procedure is general and supports other types of semantic information. We exploit this capability by including named entities, produced by a pre-existing tagger, in addition to the MN elements produced by the taggers described here. The resulting system significantly outperformed a linguistically naive baseline model (Hiero), and reached the highest scores yet reported on the NIST 2009 Urdu–English test set. This finding supports the hypothesis that both syntactic and semantic information can improve translation quality. Kathrin Baker, Michael Bloodgood, Bonnie J. Dorr, Chris Callison-Burch, Nathaniel Wesley Filardo, Christine D. Piatko, Lori S. Levin |
Comput. Linguistics | 3 |
| 2012 | Rapid understanding of scientific paper collections: Integrating statistics, text analytics, and visualizationabstractKeeping up with rapidly growing research fields, especially when there are multiple interdisciplinary sources, requires substantial effort for researchers, program managers, or venture capital investors. Current theories and tools are directed at finding a paper or website, not gaining an understanding of the key papers, authors, controversies, and hypotheses. This report presents an effort to integrate statistics, text analytics, and visualization in a multiple coordinated window environment that supports exploration. Our prototype system, Action Science Explorer (ASE), provides an environment for demonstrating principles of coordination and conducting iterative usability tests of them with interested and knowledgeable users. We developed an understanding of the value of reference management, statistics, citation text extraction, natural language summarization for single and multiple documents, filters to interactively select key papers, and network visualization to see citation patterns and identify clusters. A three‐phase usability study guided our revisions to ASE and led us to improve the testing methods. Cody Dunne, Ben Shneiderman, Robert Gove, Judith L. Klavans, Bonnie J. Dorr |
J. Assoc. Inf. Sci. Technol. | 5 |
| 2011 | Evaluating visual and statistical exploration of scientific literature networksabstractAction Science Explorer (ASE) is a tool designed to support users in rapidly generating readily consumable summaries of academic literature. It uses citation network visualization, ranking and filtering papers by network statistics, and automatic clustering and summarization techniques. We describe how early formative evaluations of ASE led to a mature system evaluation, consisting of an in-depth empirical evaluation with four domain experts. The evaluation tasks were of two types: predefined tasks to test system performance in common scenarios, and user-defined tasks to test the system's usefulness for custom exploration goals. The primary contribution of this paper is a validation of the ASE design and recommendations to provide: easy-to-understand metrics for ranking and filtering documents, user control over which document sets to explore, and overviews of the document set in coordinated views along with details-on-demand of specific papers. We contribute a taxonomy of features for literature search and exploration tools and describe exploration goals identified by our participants. Robert Gove, Cody Dunne, Ben Shneiderman, Judith L. Klavans, Bonnie J. Dorr |
VL/HCC | 5 |
| 2010 | A Modality Lexicon and its use in Automatic Tagging
Kathrin Baker, Michael Bloodgood, Bonnie J. Dorr, Nathaniel Wesley Filardo, Lori S. Levin, Christine D. Piatko |
LREC | 3 |
| 2010 | Putting the User in the Loop: Interactive Maximal Marginal Relevance for Query-Focused Summarization
Jimmy Lin, Nitin Madnani, Bonnie J. Dorr |
HLT-NAACL | 3 |
| 2010 | Generating Phrasal and Sentential Paraphrases: A Survey of Data-Driven MethodsabstractThe task of paraphrasing is inherently familiar to speakers of all languages. Moreover, the task of automatically generating or extracting semantic equivalences for the various units of language—words, phrases, and sentences—is an important part of natural language processing (NLP) and is being increasingly employed to improve the performance of several NLP applications. In this article, we attempt to conduct a comprehensive and application-independent survey of data-driven phrasal and sentential paraphrase generation methods, while also conveying an appreciation for the importance and potential use of paraphrases in the field of NLP research. Recent work done in manual and automatic construction of paraphrase corpora is also examined. We also discuss the strategies used for evaluating paraphrase generation techniques and briefly explore some future trends in paraphrase generation. Nitin Madnani, Bonnie J. Dorr |
Comput. Linguistics | 2 |
| 2010 | Interlingual annotation of parallel text corpora: a new framework for annotation and evaluationabstractAbstract This paper focuses on an important step in the creation of a system of meaning representation and the development of semantically annotated parallel corpora, for use in applications such as machine translation, question answering, text summarization, and information retrieval. The work described below constitutes the first effort of any kind to annotate multiple translations of foreign-language texts with interlingual content. Three levels of representation are introduced: deep syntactic dependencies (IL0), intermediate semantic representations (IL1), and a normalized representation that unifies conversives, nonliteral language, and paraphrase (IL2). The resulting annotated, multilingually induced, parallel corpora will be useful as an empirical basis for a wide range of research, including the development and evaluation of interlingual NLP systems and paraphrase-extraction systems as well as a host of other research and development efforts in theoretical and applied linguistics, foreign language pedagogy, translation studies, and other related disciplines. Bonnie J. Dorr, Rebecca J. Passonneau, David Farwell, Rebecca Green, Nizar Habash, Stephen Helmreich, Eduard H. Hovy, Lori S. Levin, Keith J. Miller, Teruko Mitamura, Owen Rambow, Advaith Siddharthan |
Nat. Lang. Eng. | 1 |
| 2009 | Generating High-Coverage Semantic Orientation Lexicons From Overtly Marked Words and a Thesaurus
Saif M. Mohammad, Cody Dunne, Bonnie J. Dorr |
EMNLP | 3 |
| 2009 | Using Citations to Generate surveys of Scientific Paradigms
Saif M. Mohammad, Bonnie J. Dorr, Melissa Egan, Ahmed Awadallah 0001, Pradeep Muthukrishnan, Vahed Qazvinian, Dragomir R. Radev, David M. Zajic |
HLT-NAACL | 2 |
| 2009 | Symbolic-to-statistical hybridization: extending generation-heavy machine translationabstractThe last few years have witnessed an increasing interest in hybridizing surface-based statistical approaches and rule-based symbolic approaches to machine translation (MT). Much of that work is focused on extending statistical MT systems with symbolic knowledge and components. In the brand of hybridization discussed here, we go in the opposite direction: adding statistical bilingual components to a symbolic system. Our base system is Generation-heavy machine translation (GHMT), a primarily symbolic asymmetrical approach that addresses the issue of Interlingual MT resource poverty in source-poor/target-rich language pairs by exploiting symbolic and statistical target-language resources. GHMT’s statistical components are limited to target-language models, which arguably makes it a simple form of a hybrid system . We extend the hybrid nature of GHMT by adding statistical bilingual components. We also describe the details of retargeting it to Arabic–English MT. The morphological richness of Arabic brings several challenges to the hybridization task. We conduct an extensive evaluation of multiple system variants. Our evaluation shows that this new variant of GHMT—a primarily symbolic system extended with monolingual and bilingual statistical components—has a higher degree of grammaticality than a phrase-based statistical MT system, where grammaticality is measured in terms of correct verb-argument realization and long-distance dependency translation. Nizar Habash, Bonnie J. Dorr, Christof Monz |
Mach. Transl. | 2 |
| 2009 | TER-Plus: paraphrase, semantic, and alignment enhancements to Translation Edit Rate
Matthew G. Snover, Nitin Madnani, Bonnie J. Dorr, Richard M. Schwartz |
Mach. Transl. | 3 |
| 2008 | Computing Word-Pair Antonymy
Saif M. Mohammad, Bonnie J. Dorr, Graeme Hirst |
EMNLP | 2 |
| 2008 | Language and Translation Model Adaptation using Comparable Corpora
Matthew G. Snover, Bonnie J. Dorr, Richard M. Schwartz |
EMNLP | 2 |
| 2008 | The ACL Anthology Reference Corpus: A Reference Dataset for Bibliographic Research in Computational Linguistics
Steven Bird, Robert Dale, Bonnie J. Dorr, Bryan R. Gibson, Mark Thomas Joseph, Min-Yen Kan, Dongwon Lee 0001, Brett Powley, Dragomir R. Radev, Yee Fan Tan |
LREC | 3 |
| 2008 | Single-document and multi-document summarization techniques for email threads using sentence compression
David M. Zajic, Bonnie J. Dorr, Jimmy Lin |
Inf. Process. Manag. | 2 |
| 2007 | Combining Outputs from Multiple Machine Translation Systems
Antti-Veikko I. Rosti, Necip Fazil Ayan, Bing Xiang, Spyridon Matsoukas, Richard M. Schwartz, Bonnie J. Dorr |
HLT-NAACL | 6 |
| 2007 | Exploiting aspectual features and connecting words for summarization-inspired temporal-relation extraction
Bonnie J. Dorr, Terry Gaasterland |
Inf. Process. Manag. | 1 |
| 2007 | Task-based evaluation of text summarization using Relevance Prediction
Stacy Hobson, Bonnie J. Dorr, Christof Monz, Richard M. Schwartz |
Inf. Process. Manag. | 2 |
| 2007 | Multi-candidate reduction: Sentence compression as a tool for document summarization tasks
David M. Zajic, Bonnie J. Dorr, Jimmy Lin, Richard M. Schwartz |
Inf. Process. Manag. | 2 |
| 2006 | Going Beyond AER: An Extensive Analysis of Word Alignments and Their Impact on MTabstractThis paper presents an extensive evaluation of five different alignments and investigates their impact on the corresponding MT system output. We introduce new measures for intrinsic evaluations and examine the distribution of phrases and untranslated words during decoding to identify which characteristics of different alignments affect translation. We show that precision-oriented alignments yield better MT output (translating more words and using longer phrases) than recall-oriented alignments. Necip Fazil Ayan, Bonnie J. Dorr |
ACL | 2 |
| 2006 | PCFGs with Syntactic and Prosodic Indicators of Speech RepairsabstractA grammatical method of combining two kinds of speech repair cues is presented. One cue, prosodic disjuncture, is detected by a decision tree-based ensemble classifier that uses acoustic cues to identify where normal prosody seems to be interrupted (Lickley, 1996). The other cue, syntactic parallelism, codifies the expectation that repairs continue a syntactic category that was left unfinished in the reparandum (Levelt, 1983). The two cues are combined in a Treebank PCFG whose states are split using a few simple tree transformations. Parsing performance on the Switchboard and Fisher corpora suggests that these two cues help to locate speech repairs in a synergistic way. John Hale, Izhak Shafran, Lisa Yung, Bonnie J. Dorr, Mary P. Harper, Anna Krasnyanskaya, Matthew Lease, Yang Liu 0004, Brian Roark, Matthew G. Snover, Robin Stewart |
ACL | 4 |
| 2006 | Leveraging Reusability: Cost-Effective Lexical Acquisition for Large-Scale Ontology TranslationabstractThesauri and ontologies provide important value in facilitating access to digital archives by representing underlying principles of organization. Translation of such resources into multiple languages is an important component for providing multilingual access. However, the specificity of vocabulary terms in most ontologies precludes fully-automated machine translation using general-domain lexical resources. In this paper, we present an efficient process for leveraging human translations when constructing domain-specific lexical resources. We evaluate the effectiveness of this process by producing a probabilistic phrase dictionary and translating a thesaurus of 56,000 concepts used to catalogue a large archive of oral histories. Our experiments demonstrate a cost-effective technique for accurate machine translation of large ontologies. G. Craig Murray, Bonnie J. Dorr, Jimmy Lin, Jan Hajic 0001, Pavel Pecina |
ACL | 2 |
| 2006 | Leveraging Recurrent Phrase Structure in Large-scale Ontology Translation
G. Craig Murray, Bonnie J. Dorr, Jimmy Lin, Jan Hajic 0001, Pavel Pecina |
EAMT | 2 |
| 2006 | Reranking for Sentence Boundary Detection in Conversational SpeechabstractWe present a reranking approach to sentence-like unit (SU) boundary detection, one of the EARS metadata extraction tasks. Techniques for generating relatively small n-best lists with high oracle accuracy are presented. For each candidate, features are derived from a range of information sources, including the output of a number of parsers. Our approach yields significant improvements over the best performing system from the NIST RT-04F community evaluation Brian Roark, Yang Liu 0004, Mary P. Harper, Robin Stewart, Matthew Lease, Matthew G. Snover, Izhak Shafran, Bonnie J. Dorr, John Hale, Anna Krasnyanskaya, Lisa Yung |
ICASSP (1) | 8 |
| 2006 | Parallel Syntactic Annotation of Multiple Languages
Owen Rambow, Bonnie J. Dorr, David Farwell, Rebecca Green, Nizar Habash, Stephen Helmreich, Eduard H. Hovy, Lori S. Levin, Keith J. Miller, Teruko Mitamura, Flo Reeder, Advaith Siddharthan |
LREC | 2 |
| 2006 | SParseval: Evaluation Metrics for Parsing Speech
Brian Roark, Mary P. Harper, Eugene Charniak, Bonnie J. Dorr, Mark Johnson 0001, Jeremy G. Kahn, Yang Liu 0004, Mari Ostendorf, John Hale, Anna Krasnyanskaya, Matthew Lease, Izhak Shafran, Matthew G. Snover, Robin Stewart, Lisa Yung |
LREC | 4 |
| 2006 | A Maximum Entropy Approach to Combining Word Alignments
Necip Fazil Ayan, Bonnie J. Dorr |
HLT-NAACL | 2 |
| 2006 | Automatic identification of confusable drug names
Grzegorz Kondrak, Bonnie J. Dorr |
Artif. Intell. Medicine | 2 |
| 2005 | Iterative translation disambiguation for cross-language information retrievalabstractFinding a proper distribution of translation probabilities is one of the most important factors impacting the effectiveness of a cross-language information retrieval system. In this paper we present a new approach that computes translation probabilities for a given query by using only a bilingual dictionary and a monolingual corpus in the target language. The algorithm combines term association measures with an iterative machine learning approach based on expectation maximization. Our approach considers only pairs of translation candidates and is therefore less sensitive to data-sparseness issues than approaches using higher n-grams. The learned translation probabilities are used as query term weights and integrated into a vector-space retrieval system. Results for English-German cross-lingual retrieval show substantial improvements over a baseline using dictionary lookup without term weighting. Christof Monz, Bonnie J. Dorr |
SIGIR | 2 |
| 2004 | Inducing Frame Semantic Verb Classes from WordNet and LDOCEabstractThis paper presents SemFrame, a system that induces frame semantic verb classes from WordNet and LDOCE. Semantic frames are thought to have significant potential in resolving the paraphrase problem challenging many language-based applications.When compared to the handcrafted FrameNet, SemFrame achieves its best recall-precision balance with 83.2% recall (based on SemFrame's coverage of FrameNet frames) and 73.8% precision (based on SemFrame verbs' semantic relatedness to frame-evoking verbs). The next best performing semantic verb classes achieve 56.9% recall and 55.0% precision. Rebecca Green, Bonnie J. Dorr, Philip Resnik |
ACL | 2 |
| 2004 | Identification of Confusable Drug Names: A New Approach and Evaluation Methodology
Grzegorz Kondrak, Bonnie J. Dorr |
COLING | 2 |
| 2003 | Acquisition of bilingual MT lexicons from OCRed dictionariesabstractThis paper describes an approach to analyzing the lexical structure of OCRed bilingual dictionaries to construct resources suited for machine translation of low-density languages, where online resources are limited. A rule-based, an HMM-based, and a post-processed HMM-based method are used for rapid construction of MT lexicons based on systematic structural clues provided in the original dictionary. We evaluate the effectiveness of our techniques, concluding that: (1) the rule-based method performs better with dictionaries where the font is not an important distinguishing feature for determining information types; (2) the post-processed stochastic method improves the results of the stochastic method for phrasal entries; and (3) Our resulting bilingual lexicons are comprehensive enough to provide the basis for reasonable translation results when compared to human translations. Burcu Karagol Ayan, David S. Doermann, Bonnie J. Dorr |
MTSummit | 3 |
| 2003 | A Categorial Variation Database for English
Nizar Habash, Bonnie J. Dorr |
HLT-NAACL | 2 |
| 2003 | Desparately Seeking Cebuano
Douglas W. Oard, David S. Doermann, Bonnie J. Dorr, Daqing He, Philip Resnik, Amy Weinberg, William J. Byrne, Sanjeev Khudanpur, David Yarowsky, Anton Leuski, Philipp Koehn, Kevin Knight |
HLT-NAACL | 3 |
| 2003 | Hybrid Natural Language Generation from Lexical Conceptual Structures
Nizar Habash, Bonnie J. Dorr, David R. Traum |
Mach. Transl. | 2 |
| 2003 | Rapid porting of DUSTer to HindiabstractThe frequent occurrence of divergences —structural differences between languages---presents a great challenge for statistical word-level alignment and machine translation. This paper describes the adaptation of DUSTer, a divergence unraveling package, to Hindi during the DARPA TIDES-2003 Surprise Language Exercise. We show that it is possible to port DUSTer to Hindi in under 3 days. Bonnie J. Dorr, Necip Fazil Ayan, Nizar Habash, Nitin Madnani, Rebecca Hwa |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2003 | Cross-language headline generation for HindiabstractThis paper presents new approaches to headline generation for English newspaper texts, with an eye toward the production of document surrogates for document selection in cross-language information retrieval. This task is difficult because the user must make decisions about relevance based on (often poor) translations of retrieved documents. To facilitate the decision-making process we need translations that can be assessed rapidly and accurately; our approach is to provide an English headline for the non-English document. We describe two approaches to headline generation and their application to the recent DARPA TIDES-2003 Surprise Language Exercise for Hindi. For comparison, we also implemented an alternative method for surrogate generation: a system that produces topic lists for (Hindi) articles. We present the results of a series of experiments comparing each of these approaches. We demonstrate in both automatic and human evaluations that our linguistically motivated approach outperforms two other surrogate-generation methods: a statistical system and a topic discovery system. Bonnie J. Dorr, David M. Zajic, Richard M. Schwartz |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2002 | Construction of a Chinese-English Verb Lexicon for Machine Translation and Embedded Multilingual Applications
Bonnie J. Dorr, Gina-Anne Levow, Dekang Lin |
Mach. Transl. | 1 |
| 2001 | Mapping Lexical Entries in a Verbs Database to WordNet SensesabstractThis paper describes automatic techniques for mapping 9611 entries in a database of English verbs to WordNet senses. The verbs were initially grouped into 491 classes based on syntactic features. Mapping these verbs into WordNet senses provides a resource that supports disambiguation in multilingual applications such as machine translation and cross-language information retrieval. Our techniques make use of (1) a training set of 1791 disambiguated entries, representing 1442 verb entries from 167 classes; (2) word sense probabilities, from frequency counts in a tagged corpus; (3) semantic similarity of WordNet senses for verbs within the same class; (4) probabilistic correlations between WordNet data and attributes of the verb classes. The best results achieved 72% precision and 58% recall, versus a lower bound of 62% precision and 38% recall for assigning the most frequently occurring WordNet sense, and an upper bound of 87% precision and 75% recall for human judgment. Rebecca Green, Lisa Pearl, Bonnie J. Dorr, Philip Resnik |
ACL | 3 |
| 2001 | Interpretation of Compound Nominals Using WordNet
Leslie Barrett, Anthony R. Davis, Bonnie J. Dorr |
CICLing | 3 |
| 2001 | Large scale language independent generation using thematic hierarchies
Nizar Habash, Bonnie J. Dorr |
MTSummit | 2 |
| 2001 | Review of Natural Language Processing in R.A. Wilson and F.C. Keil (Eds.), The MIT Encyclopedia of the Cognitive Sciences
Bonnie J. Dorr |
Artif. Intell. | 1 |
| 2000 | Chinese-English Semantic Resource Construction
Bonnie J. Dorr, Gina-Anne Levow, Dekang Lin, Scott C. Thomas |
LREC | 1 |
| 1998 | Evaluation of Eurowordnet- and LCS- based lexical resources for machine translation
Bonnie J. Dorr, Maria Antònia Martí, Irene Castellón |
LREC | 1 |
| 1998 | Evaluating resources for query translation in cross-language information retrieval
Bonnie J. Dorr, Douglas W. Oard |
LREC | 1 |
| 1997 | Deriving Verbal and Compositional Lexical Aspect for NLP ApplicationsabstractVerbal and compositional lexical aspect provide the underlying temporal structure of events. Knowledge of lexical aspect, e.g., (a)telicity, is therefore required for interpreting event sequences in discourse (Dowty, 1986; Moens and Steedman, 1988; Passoneau, 1988), interfacing to temporal databases (Androutsopoulos, 1996), processing temporal modifiers (Antonisse, 1994), describing allowable alternations and their semantic effects (Resnik, 1996; Tenny, 1994), and selecting tense and lexical items for natural language generation ((Dorr and Olsen, 1996; Klavans and Chodorow, 1992), cf. (Slobin and Bocaz, 1988)). We show that it is possible to represent lexical aspect---both verbal and compositional---on a large scale, using Lexical Conceptual Structure (LCS) representations of verbs in the classes cataloged by Levin (1993). We show how proper consideration of these universal pieces of verb meaning may be used to refine lexical representations and derive a range of meanings from combinations of LCS representations. A single algorithm may therefore be used to determine lexical aspect classes and features at both verbal and sentence levels. Finally, we illustrate how knowledge of lexical aspect facilitates the interpretation of events in NLP applications. Bonnie J. Dorr, Mari Broman Olsen |
ACL | 1 |
| 1997 | Large-Scale Dictionary Construction for Foreign Language Tutoring and Interlingual Machine Translation
Bonnie J. Dorr |
Mach. Transl. | 1 |
| 1996 | Role of Word Sense Disambiguation in Lexical Acquisition: Predicting Semantics from Syntactic Cues
Bonnie J. Dorr, Douglas A. Jones |
COLING | 1 |
| 1996 | Multilingual generation: The role of telicity in lexical choice and syntactic realization
Bonnie J. Dorr, Mari Broman Olsen |
Mach. Transl. | 1 |
| 1995 | Selecting Tense, Aspect, and Connecting Words In Language Generation
Bonnie J. Dorr, Terry Gaasterland |
IJCAI | 1 |
| 1995 | Efficient Parsing for Korean and English: A Parameterized Message- Passing Approach
Bonnie J. Dorr, Jye-hoon Lee, Dekang Lin, Sungki Suh |
Comput. Linguistics | 1 |
| 1995 | Introduction: Special issue on building lexicons for machine translation
Bonnie J. Dorr, Judith L. Klavans |
Mach. Transl. | 1 |
| 1995 | Toward a lexicalized grammar for interlinguas
Clare R. Voss, Bonnie J. Dorr |
Mach. Transl. | 2 |
| 1994 | Query Transformation Techniques for Interoperable Query Processing in Cooperative Information Systems
Louiqa Raschid, Yahui Chang, Bonnie J. Dorr |
CoopIS | 3 |
| 1994 | Transforming Queries from a Relational Schema to an Equivalent Object Schema: A Prototype Based on F-logic
Yahui Chang, Louiqa Raschid, Bonnie J. Dorr |
ISMIS | 3 |
| 1994 | Machine Translation Divergences: A Formal Description and Proposed Solution
Bonnie J. Dorr |
Comput. Linguistics | 1 |
| 1994 | From syntactic encodings to thematic roles: Building lexical entries for interlingual MT
Bonnie J. Dorr, Joseph Garman, Amy Weinberg |
Mach. Transl. | 1 |
| 1994 | Introduction: Special issue on building lexicons for machine translation
Bonnie J. Dorr, Judith L. Klavans |
Mach. Transl. | 1 |
| 1993 | Machine Translation of Spatial Expressions: Defining the Relation between an Interlingua and a Knowledge Representation System
Bonnie J. Dorr, Clare R. Voss |
AAAI | 1 |
| 1993 | Interoperable Query Processing with Multiple Heterogeneous Knowledge ServersabstractThis paper describes a technique for information mediation when multiple heterogeneous knowledge and data servers are to be accessed during query processing. One problem is building an intelligent interface between each knowledge server (KS) and its processor (KP); and the second is to provide interoperability among multiple KP/KS so that a query may be answered using information from multiple sources. We present example scenarios which highlight these problems and then outline query mapping and transformation techniques that are applicable. The techniques for solving the interoperability problems involve representations in some canonical form. This includes a canonical representation (CR) corresponding to each KP/KS pair and a merged CR (MCR) to represent the mapping among the CRs. The MCR and CRs include relevant information obtained from a source query, and heterogeneous mapping (het-map) information, for all possible mappings among the multiple servers. The knowledge in the canonical form must be represented so that it can be easily during query transformation. Louiqa Raschid, Yahui Chang, Bonnie J. Dorr |
CIKM | 3 |
| 1993 | Interlingual Machine Translation: A Parameterized Approach
Bonnie J. Dorr |
Artif. Intell. | 1 |
| 1993 | A first-pass approach for evaluating machine translation systems
Pamela W. Jordan, Bonnie J. Dorr, John W. Benoit |
Mach. Transl. | 2 |
| 1992 | A Parameterized Approach to Integrating Aspect with Lexical-Semanics for Machine TranslationabstractThis paper discusses how a two-level knowledge representation model for machine translation integrates aspectual information with lexical-semantic information by means of parameterization. The integration of aspect with lexical-semantics is especially critical in machine translation because of the lexical selection and aspectual realization processes that operate during the production of the target-language sentence: there are often a large number of lexical and aspectual possibilities to choose from in the production of a sentence from a lexical semantic representation. Aspectual information from the source-language sentence constrains the choice of target-language terms. In turn, the target-language terms limit the possibilities for generation of aspect. Thus, there is a two-way communication channel between the two processes. This paper will show that the selection/realization processes may be parameterized so that they operate uniformly across more than one language and it will describe how the parameter-based approach is currently being used as the basis for extraction of aspectual information from corpora. Bonnie J. Dorr |
ACL | 1 |
| 1992 | Parameterization of the Interlingua in Machine Translation
Bonnie J. Dorr |
COLING | 1 |
| 1992 | The use of lexical semantics in interlingual machine translation
Bonnie J. Dorr |
Mach. Transl. | 1 |
| 1990 | Solving Thematic Divergences in Machine TranslationabstractThough most translation systems have some mechanism for translating certain types of divergent predicate-argument structures, they do not provide a general procedure that takes advantage of the relationship between lexical-semantic structure and syntactic structure. A divergent predicate-argument structure is one in which the predicate (e.g., the main verb) or its arguments (e.g., the subject and object) do not have the same syntactic ordering properties for both the source and target language. To account for such ordering differences, a machine translator must consider language-specific syntactic idiosyncrasies that distinguish a target language from a source language, while making use of lexical-semantic uniformities that tie the two languages together. This paper describes the mechanisms used by the UNITRAN machine translation system for mapping an underlying lexical-conceptual structure to a syntactic structure (and vice versa), and it shows how these mechanisms coupled with a set of general linking routines solve the problem of thematic divergence in machine translation. Bonnie J. Dorr |
ACL | 1 |
| 1987 | UNITRAN: An Interlingual Approach to Machine Translation
Bonnie J. Dorr |
AAAI | 1 |