William B. Dolan

dblp:13/486 · also Bill Dolan · DBLP profile ↗
← Back
48ranked-venue papers
2as first author
11since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2024 GENEVA: GENErating and Visualizing branching narratives using LLMs
abstract
Dialogue-based Role Playing Games (RPGs) require powerful storytelling. The narratives of these may take years to write and typically involve a large creative team. In this work, we demonstrate the potential of large generative text models to assist this process. GENEVA, a prototype tool, generates a rich narrative graph with branching and reconverging storylines that match a high-level narrative description and constraints provided by the designer. A large language model (LLM), GPT-4, is used to generate the branching narrative and to render it in a graph format in a two-step process. We illustrate the use of GENEVA in generating new branching narratives for four well-known stories under different contextual constraints. This tool has the potential to assist in game development, simulations, and other applications with game-like properties.
Jorge Leandro, Sudha Rao, Michael Xu, Weijia Xu, Nebojsa Jojic, Chris Brockett, William B. Dolan
CoG7
2024 Player-Driven Emergence in LLM-Driven Game Narrative
abstract
We explore how interaction with large language models (LLMs) can give rise to emergent behaviors, empowering players to participate in the evolution of game narratives. Our testbed is a text-adventure game in which players attempt to solve a mystery under a fixed narrative premise, but can freely interact with non-player characters generated by GPT-4, a large language model. We recruit 28 gamers to play the game and use GPT-4 to automatically convert the game logs into a node-graph representing the narrative in the player’s gameplay. We find that through their interactions with the non-deterministic behavior of the LLM, players are able to discover interesting new emergent nodes that were not a part of the original narrative but have potential for being fun and engaging. Players that created the most emergent nodes tended to be those that often enjoy games that facilitate discovery, exploration and experimentation.
Jessica Quaye, Sudha Rao, Weijia Xu, Portia Botchway, Chris Brockett, Nebojsa Jojic, Gabriel DesGarennes, Ken Lobb, Michael Xu, Jorge Leandro, Claire Jin, William B. Dolan
CoG13
2024 Investigating Agency of LLMs in Human-AI Collaboration Tasks
abstract
Ashish Sharma, Sudha Rao, Chris Brockett, Akanksha Malhotra, Nebojsa Jojic, Bill Dolan. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Sudha Rao, Chris Brockett, Akanksha Malhotra, Nebojsa Jojic, William B. Dolan
EACL (1)6
2023 Towards More Efficient Insertion Transformer with Fractional Positional Encoding
abstract
Auto-regressive neural sequence models have been shown to be effective across text generation tasks.However, their left-to-right decoding order prevents generation from being parallelized.Insertion Transformer (Stern et al., 2019) is an attractive alternative that allows outputting multiple tokens in a single generation step.Nevertheless, due to the incompatibility between absolute positional encoding and insertion-based generation schemes, it needs to refresh the encoding of every token in the generated partial hypothesis at each step, which could be costly.We design a novel reusable positional encoding scheme for Insertion Transformers called Fractional Positional Encoding (FPE), which allows reusing representations calculated in previous steps.Empirical studies on various text generation tasks demonstrate the effectiveness of FPE, which leads to floating-point operation reduction and latency improvements on batched decoding.
Zhisong Zhang, Yizhe Zhang 0002, William B. Dolan
EACL3
2023 Interactive Text Generation
abstract
Felix Faltings, Michel Galley, Kianté Brantley, Baolin Peng, Weixin Cai, Yizhe Zhang, Jianfeng Gao, Bill Dolan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Felix Faltings, Michel Galley, Kianté Brantley, Baolin Peng, Weixin Cai, Yizhe Zhang 0002, Jianfeng Gao 0001, William B. Dolan
EMNLP8
2022 RetGen: A Joint Framework for Retrieval and Grounded Text Generation Modeling
abstract
Recent advances in large-scale pre-training such as GPT-3 allow seemingly high quality text to be generated from a given prompt. However, such generation systems often suffer from problems of hallucinated facts, and are not inherently designed to incorporate useful external information. Grounded generation models appear to offer remedies, but their training typically relies on rarely-available parallel data where information-relevant documents are provided for context. We propose a framework that alleviates this data constraint by jointly training a grounded generator and document retriever on the language model signal. The model learns to reward retrieval of the documents with the highest utility in generation, and attentively combines them using a Mixture-of-Experts (MoE) ensemble to generate follow-on text. We demonstrate that both generator and retriever can take advantage of this joint training and work synergistically to produce more informative and relevant text in both prose and dialogue generation.
Yizhe Zhang 0002, Xiang Gao 0011, Yuwei Fang, Chris Brockett, Michel Galley, Jianfeng Gao 0001, William B. Dolan
AAAI8
2022 A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text Generation
abstract
Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, Bill Dolan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Tianyu Liu 0001, Yizhe Zhang 0002, Chris Brockett, Zhifang Sui, Weizhu Chen, William B. Dolan
ACL (1)7
2021 A Controllable Model of Grounded Response Generation
abstract
Current end-to-end neural conversation models inherently lack the flexibility to impose semantic control in the response generation process, often resulting in uninteresting responses. Attempts to boost informativeness alone come at the expense of factual accuracy, as attested by pretrained language models' propensity to "hallucinate" facts. While this may be mitigated by access to background knowledge, there is scant guarantee of relevance and informativeness in generated responses. We propose a framework that we call controllable grounded response generation (CGRG), in which lexical control phrases are either provided by a user or automatically extracted by a control phrase predictor from dialogue context and grounding knowledge. Quantitative and qualitative results show that, using this framework, a transformer based model with a novel inductive attention mechanism, trained on a conversation-like Reddit dataset, outperforms strong generation baselines.
Zeqiu Wu, Michel Galley, Chris Brockett, Yizhe Zhang 0002, Xiang Gao 0011, Chris Quirk, Rik Koncel-Kedziorski, Jianfeng Gao 0001, Hannaneh Hajishirzi, Mari Ostendorf, William B. Dolan
AAAI11
2021 Contrastive Multi-document Question Generation
abstract
Woon Sang Cho, Yizhe Zhang, Sudha Rao, Asli Celikyilmaz, Chenyan Xiong, Jianfeng Gao, Mengdi Wang, Bill Dolan. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Woon Sang Cho, Yizhe Zhang 0002, Sudha Rao, Asli Celikyilmaz, Chenyan Xiong, Jianfeng Gao 0001, Mengdi Wang 0001, William B. Dolan
EACL8
2021 Text Editing by Command
abstract
Felix Faltings, Michel Galley, Gerold Hintz, Chris Brockett, Chris Quirk, Jianfeng Gao, Bill Dolan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Felix Faltings, Michel Galley, Gerold Hintz, Chris Brockett, Chris Quirk, Jianfeng Gao 0001, William B. Dolan
NAACL-HLT7
2021 Contextualized Perturbation for Textual Adversarial Attack
abstract
Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, Bill Dolan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Dianqi Li, Yizhe Zhang 0002, Hao Peng 0009, Liqun Chen 0001, Chris Brockett, Ming-Ting Sun, William B. Dolan
NAACL-HLT7
2020 A Recipe for Creating Multimodal Aligned Datasets for Sequential Tasks
abstract
Many high-level procedural tasks can be decomposed into sequences of instructions that vary in their order and choice of tools. In the cooking domain, the web offers many partially-overlapping text and video recipes (i.e. procedures) that describe how to make the same dish (i.e. high-level task). Aligning instructions for the same dish across different sources can yield descriptive visual explanations that are far richer semantically than conventional textual instructions, providing commonsense insight into how real-world procedures are structured. Learning to align these different instruction sets is challenging because: a) different recipes vary in their order of instructions and use of ingredients; and b) video instructions can be noisy and tend to contain far more information than text instructions. To address these challenges, we first use an unsupervised alignment algorithm that learns pairwise alignments between instructions of different recipes for the same dish. We then use a graph algorithm to derive a joint alignment between multiple text and multiple video recipes for the same dish. We release the Microsoft Research Multimodal Aligned Recipe Corpus containing 150K pairwise alignments between recipes across 4,262 dishes with rich commonsense information.
Angela S. Lin, Sudha Rao, Asli Celikyilmaz, Elnaz Nouri, Chris Brockett, Debadeepta Dey, William B. Dolan
ACL7
2020 Dialogue Response Ranking Training with Large-Scale Human Feedback Data
abstract
Existing open-domain dialog models are generally trained to minimize the perplexity of target human responses.However, some human replies are more engaging than others, spawning more followup interactions.Current conversational models are increasingly capable of producing turns that are context-relevant, but in order to produce compelling agents, these models need to be able to predict and optimize for turns that are genuinely engaging.We leverage social media feedback data (number of replies and upvotes) to build a large-scale training dataset for feedback prediction.To alleviate possible distortion between the feedback and engagingness, we convert the ranking problem to a comparison of response pairs which involve few confounding factors.We trained DIALOGRPT, a set of GPT-2 based models on 133M pairs of human feedback data and the resulting ranker outperformed several baselines.Particularly, our ranker outperforms the conventional dialog perplexity baseline with a large margin on predicting Reddit feedback.We finally combine the feedback prediction models and a human-like scoring model to rank the machine-generated dialog responses.Crowd-sourced human evaluation shows that our ranking method correlates better with real human preferences than baseline models. 1
Xiang Gao 0011, Yizhe Zhang 0002, Michel Galley, Chris Brockett, William B. Dolan
EMNLP (1)5
2020 Substance over Style: Document-Level Targeted Content Transfer
abstract
Existing language models excel at writing from scratch, but many real-world scenarios require rewriting an existing document to fit a set of constraints.Although sentence-level rewriting has been fairly well-studied, little work has addressed the challenge of rewriting an entire document coherently.In this work, we introduce the task of document-level targeted content transfer and address it in the recipe domain, with a recipe as the document and a dietary restriction (such as vegan or dairy-free) as the targeted constraint.We propose a novel model for this task based on the generative pretrained language model (GPT-2) and train on a large number of roughly-aligned recipe pairs. 1 Both automatic and human evaluations show that our model out-performs existing methods by generating coherent and diverse rewrites that obey the constraint while remaining close to the original document.Finally, we analyze our model's rewrites to assess progress toward the goal of making language generation more attuned to constraints that are substantive rather than stylistic.
Allison Hegel, Sudha Rao, Asli Celikyilmaz, William B. Dolan
EMNLP (1)4
2020 POINTER: Constrained Progressive Text Generation via Insertion-based Generative Pre-training
abstract
Large-scale pre-trained language models, such as BERT and GPT-2, have achieved excellent performance in language representation learning and free-form text generation.However, these models cannot be directly employed to generate text under specified lexical constraints.To address this challenge, we present POINTER 1 , a simple yet novel insertion-based approach for hard-constrained text generation.The proposed method operates by progressively inserting new tokens between existing tokens in a parallel manner.This procedure is recursively applied until a sequence is completed.The resulting coarse-to-fine hierarchy makes the generation process intuitive and interpretable.We pre-train our model with the proposed progressive insertion-based objective on a 12GB Wikipedia dataset, and finetune it on downstream hard-constrained generation tasks.Non-autoregressive decoding yields an empirically logarithmic time complexity during inference time.Experimental results on both News and Yelp datasets demonstrate that POINTER achieves state-of-the-art performance on constrained text generation.We released the pre-trained models and the source code to facilitate future research 2 .
Yizhe Zhang 0002, Guoyin Wang 0002, Chunyuan Li, Zhe Gan, Chris Brockett, William B. Dolan
EMNLP (1)6
2019 Conversing by Reading: Contentful Neural Conversation with On-demand Machine Reading
abstract
Lianhui Qin, Michel Galley, Chris Brockett, Xiaodong Liu, Xiang Gao, Bill Dolan, Yejin Choi, Jianfeng Gao. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Lianhui Qin, Michel Galley, Chris Brockett, Xiaodong Liu 0003, Xiang Gao 0011, William B. Dolan, Yejin Choi 0001, Jianfeng Gao 0001
ACL (1)6
2019 Vision-Based Navigation With Language-Based Assistance via Imitation Learning With Indirect Intervention
abstract
We present Vision-based Navigation with Language-based Assistance (VNLA), a grounded vision-language task where an agent with visual perception is guided via language to find objects in photorealistic indoor environments. The task emulates a real-world scenario in that (a) the requester may not know how to navigate to the target objects and thus makes requests by only specifying high-level end-goals, and (b) the agent is capable of sensing when it is lost and querying an advisor, who is more qualified at the task, to obtain language subgoals to make progress. To model language-based assistance, we develop a general framework termed Imitation Learning with Indirect Intervention (I3L), and propose a solution that is effective on the VNLA task. Empirical results show that this approach significantly improves the success rate of the learning agent over other baselines on both seen and unseen environments. Our code and data are publicly available at https://github.com/debadeepta/vnla .
Debadeepta Dey, Chris Brockett, William B. Dolan
CVPR4
2019 Structuring Latent Spaces for Stylized Response Generation
abstract
Xiang Gao, Yizhe Zhang, Sungjin Lee, Michel Galley, Chris Brockett, Jianfeng Gao, Bill Dolan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xiang Gao 0011, Yizhe Zhang 0002, Michel Galley, Chris Brockett, Jianfeng Gao 0001, William B. Dolan
EMNLP/IJCNLP (1)7
2019 Domain Adaptive Text Style Transfer
abstract
Dianqi Li, Yizhe Zhang, Zhe Gan, Yu Cheng, Chris Brockett, Bill Dolan, Ming-Ting Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Dianqi Li, Yizhe Zhang 0002, Zhe Gan, Yu Cheng 0001, Chris Brockett, William B. Dolan, Ming-Ting Sun
EMNLP/IJCNLP (1)6
2018 A Knowledge-Grounded Neural Conversation Model
abstract
Neural network models are capable of generating extremely natural sounding conversational interactions. However, these models have been mostly applied to casual scenarios (e.g., as “chatbots”) and have yet to demonstrate they can serve in more useful conversational applications. This paper presents a novel, fully data-driven, and knowledge-grounded neural conversation model aimed at producing more contentful responses. We generalize the widely-used Sequence-to-Sequence (Seq2Seq) approach by conditioning responses on both conversation history and external “facts”, allowing the model to be versatile and applicable in an open-domain setting. Our approach yields significant improvements over a competitive Seq2Seq baseline. Human judges found that our outputs are significantly more informative.
Marjan Ghazvininejad, Chris Brockett, Ming-Wei Chang, William B. Dolan, Jianfeng Gao 0001, Scott Yih, Michel Galley
AAAI4
2018 Emotional Dialogue Generation using Image-Grounded Language Models
abstract
Computer-based conversational agents are becoming ubiquitous. However, for these systems to be engaging and valuable to the user, they must be able to express emotion, in addition to providing informative responses. Humans rely on much more than language during conversations; visual information is key to providing context. We present the first example of an image-grounded conversational agent using visual sentiment, facial expression and scene features. We show that key qualities of the generated dialogue can be manipulated by the features used for training the agent. We evaluate our model on a large and very challenging real-world dataset of conversations from social media (Twitter). The image-grounding leads to significantly more informative, emotional and specific responses, and the exact qualities can be tuned depending on the image features used. Furthermore, our model improves the objective quality of dialogue responses when evaluated on standard natural language metrics.
Bernd Huber, Daniel McDuff, Chris Brockett, Michel Galley, William B. Dolan
CHI5
2018 Generating More Interesting Responses in Neural Conversation Models with Distributional Constraints
abstract
Neural conversation models tend to generate safe, generic responses for most inputs.This is due to the limitations of likelihoodbased decoding objectives in generation tasks with diverse outputs, such as conversation.To address this challenge, we propose a simple yet effective approach for incorporating side information in the form of distributional constraints over the generated responses.We propose two constraints that help generate more content rich responses that are based on a model of syntax and topics (Griffiths et al., 2005) and semantic similarity (Arora et al., 2016).We evaluate our approach against a variety of competitive baselines, using both automatic metrics and human judgments, showing that our proposed approach generates responses that are much less generic without sacrificing plausibility.A working demo of our code can be found at https://github.com/abaheti95/ DC-NeuralConversation.
Ashutosh Baheti, Alan Ritter, Jiwei Li 0001, William B. Dolan
EMNLP4
2018 Generating Informative and Diverse Conversational Responses via Adversarial Information Maximization
abstract
Responses generated by neural conversational models tend to lack informativeness and diversity. We present Adversarial Information Maximization (AIM), an adversarial learning framework that addresses these two related but distinct problems. To foster response diversity, we leverage adversarial training that allows distributional matching of synthetic and real responses. To improve informativeness, our framework explicitly optimizes a variational lower bound on pairwise mutual information between query and response. Empirical results from automatic and human evaluations demonstrate that our methods significantly boost informativeness and diversity.
Yizhe Zhang 0002, Michel Galley, Jianfeng Gao 0001, Zhe Gan, Xiujun Li, Chris Brockett, William B. Dolan
NeurIPS7
2017 Multi-Task Learning for Speaker-Role Adaptation in Neural Conversation Models
abstract
Building a persona-based conversation agent is challenging owing to the lack of large amounts of speaker-specific conversation data for model training. This paper addresses the problem by proposing a multi-task learning approach to training neural conversation models that leverages both conversation data across speakers and other types of data pertaining to the speaker and speaker roles to be modeled. Experiments show that our approach leads to significant improvements over baseline model quality, generating responses that capture more precisely speakers’ traits and speaking styles. The model offers the benefits of being algorithmically simple and easy to implement, and not relying on large quantities of data representing specific individual speakers.
Yi Luan, Chris Brockett, William B. Dolan, Jianfeng Gao 0001, Michel Galley
IJCNLP(1)3
2017 Image-Grounded Conversations: Multimodal Context for Natural Question and Response Generation
abstract
The popularity of image sharing on social media and the engagement it creates between users reflect the important role that visual context plays in everyday conversations. We present a novel task, Image Grounded Conversations (IGC), in which natural-sounding conversations are generated about a shared image. To benchmark progress, we introduce a new multiple reference dataset of crowd-sourced, event-centric conversations on images. IGC falls on the continuum between chit-chat and goal-directed conversation models, where visual grounding constrains the topic of conversation to event-driven utterances. Experiments with models trained on social media data show that the combination of visual and textual context enhances the quality of generated conversational turns. In human evaluation, the gap between human performance and that of both neural and retrieval architectures suggests that multi-modal IGC presents an interesting challenge for dialog research.
Nasrin Mostafazadeh, Chris Brockett, William B. Dolan, Michel Galley, Jianfeng Gao 0001, Georgios Spithourakis, Lucy Vanderwende
IJCNLP(1)3
2016 Microsummarization of Online Reviews: An Experimental Study
abstract
Mobile and location-based social media applications provide platforms for users to share brief opinions about products, venues, and services. These quickly typed opinions, or microreviews, are a valuable source of current sentiment on a wide variety of subjects. However, there is currently little research on how to mine this information to present it back to users in easily consumable way. In this paper, we introduce the task of microsummarization, which combines sentiment analysis, summarization, and entity recognition in order to surface key content to users. We explore unsupervised and supervised methods for this task, and find we can reliably extract relevant entities and the sentiment targeted towards them using crowdsourced labels as supervision. In an end-to-end evaluation, we find our best-performing system is vastly preferred by judges over a traditional extractive summarization approach. This work motivates an entirely new approach to summarization, incorporating both sentiment analysis and item extraction for modernized, at-a-glance presentation of public opinion.
Rebecca Mason, Benjamin Gaska, Benjamin Van Durme, Pallavi Choudhury, Ted Hart, William B. Dolan, Kristina Toutanova, Margaret Mitchell
AAAI6
2016 A Persona-Based Neural Conversation Model
abstract
Jiwei Li, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao, Bill Dolan. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016.
Jiwei Li 0001, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao 0001, William B. Dolan
ACL (1)6
2016 A Diversity-Promoting Objective Function for Neural Conversation Models
abstract
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, Bill Dolan. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Jiwei Li 0001, Michel Galley, Chris Brockett, Jianfeng Gao 0001, William B. Dolan
HLT-NAACL5
2015 A Neural Network Approach to Context-Sensitive Generation of Conversational Responses
abstract
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, Bill Dolan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao 0001, William B. Dolan
HLT-NAACL9
2014 Extracting Lexically Divergent Paraphrases from Twitter
abstract
We present MultiP (Multi-instance Learning Paraphrase Model), a new model suited to identify paraphrases within the short messages on Twitter. We jointly model paraphrase relations between word and sentence pairs and assume only sentence-level annotations during learning. Using this principled latent variable model alone, we achieve the performance competitive with a state-of-the-art method which combines a latent space model with a feature-based supervised classifier. Our model also captures lexically divergent paraphrases that differ from yet complement previous methods; combining our model with previous work significantly outperforms the state-of-the-art. In addition, we present a novel annotation methodology that has allowed us to crowdsource a paraphrase corpus from Twitter. We make this new dataset available to the research community.
Wei Xu 0004, Alan Ritter, Chris Callison-Burch, William B. Dolan, Yangfeng Ji
Trans. Assoc. Comput. Linguistics4
2013 Lightly Supervised Learning of Procedural Dialog Systems
Svitlana Volkova, Pallavi Choudhury, Chris Quirk, William B. Dolan, Luke Zettlemoyer
ACL (1)4
2013 Paraphrase features to improve natural language understanding
Ruhi Sarikaya, Chris Brockett, Chris Quirk, William B. Dolan
INTERSPEECH5
2013 Learning to Relate Literal and Sentimental Descriptions of Visual Properties
Mark Yatskar, Svitlana Volkova, Asli Celikyilmaz, William B. Dolan, Luke Zettlemoyer
HLT-NAACL4
2013 Introduction to special section on paraphrasing
abstract
No abstract available.
Haifeng Wang 0001, William B. Dolan, Idan Szpektor
ACM Trans. Intell. Syst. Technol.2
2012 Paraphrasing for Style
Wei Xu 0004, Alan Ritter, William B. Dolan, Ralph Grishman, Colin Cherry
COLING3
2012 CLex: A Lexicon for Exploring Color, Concept and Emotion Associations in Language
Svitlana Volkova, William B. Dolan, Theresa Wilson
EACL2
2011 Collecting Highly Parallel Data for Paraphrase Evaluation
David L. Chen, William B. Dolan
ACL2
2011 Data-Driven Response Generation in Social Media
Alan Ritter, Colin Cherry, William B. Dolan
EMNLP3
2010 Unsupervised Modeling of Twitter Conversations
Alan Ritter, Colin Cherry, William B. Dolan
HLT-NAACL3
2010 Recognizing textual entailment: Rational, evaluation and approaches - Erratum
abstract
Due to publisher error, this article was omitted from the printed issue ofNatural Language Engineeringvolume 15 issue 4. It is published online in the correct volume ( journals.cambridge.org/nle ) and also printed here in volume 16 issue 1. Sincere apologies are extended to the authors for this error.
Ido Dagan, William B. Dolan, Bernardo Magnini, Dan Roth 0001
Nat. Lang. Eng.2
2008 Using Contextual Speller Techniques and Language Modeling for ESL Error Correction
Michael Gamon, Jianfeng Gao 0001, Chris Brockett, Alexandre Klementiev, William B. Dolan, Dmitriy Belenko, Lucy Vanderwende
IJCNLP5
2008 A Web-based English Proofing System for English as a Second Language Users
Xing Yi, Jianfeng Gao 0001, William B. Dolan
IJCNLP3
2006 Correcting ESL Errors Using Phrasal SMT Techniques
abstract
This paper presents a pilot study of the use of phrasal Statistical Machine Translation (SMT) techniques to identify and correct writing errors made by learners of English as a Second Language (ESL). Using examples of mass noun errors found in the Chinese Learner Error Corpus (CLEC) to guide creation of an engineered training set, we show that application of the SMT paradigm can capture errors not well addressed by widely-used proofing tools designed for native speakers. Our system was able to correct 61.81% of mistakes in a set of naturally-occurring examples of mass noun errors found on the World Wide Web, suggesting that efforts to collect alignable corpora of pre- and post-editing ESL writing samples offer can enable the development of SMT-based writing assistance tools capable of repairing many of the complex syntactic and lexical problems found in the writing of ESL learners.
Chris Brockett, William B. Dolan, Michael Gamon
ACL2
2004 Unsupervised Construction of Large Paraphrase Corpora: Exploiting Massively Parallel News Sources
William B. Dolan, Chris Quirk, Chris Brockett
COLING1
2004 Monolingual Machine Translation for Paraphrase Generation
Chris Quirk, Chris Brockett, William B. Dolan
EMNLP3
2001 Achieving commercial-quality translation with example-based methods
abstract
We describe MSR-MT, a large-scale example-based machine translation system under development for several language pairs. Trained on aligned English-Spanish technical prose, a blind evaluation shows that MSR-MT’s integration of rule-based parsers, example based processing, and statistical techniques produces translations whose quality in this domain exceeds that of uncustomized commercial MT systems.
Stephen D. Richardson, William B. Dolan, Arul Menezes, Jessie Pinkham
MTSummit2
1999 Less is more: Eliminating index terms from subordinate clauses
abstract
We perform a linguistic analysis of documents during indexing for information retrieval. By eliminating index terms that occur only in subordinate clauses, index size is reduced by approximately 30% without adversely affecting precision or recall. These results hold for two corpora: a sample of the world wide web and an electronic encyclopedia.
Simon Corston-Oliver, William B. Dolan
ACL2
1994 Word Sense Ambiguation: Clustering Related Senses
William B. Dolan
COLING1