VLDB 2026 Research / reviewers in the wild / expert
Chris Brockett
dblp:95/6500
· DBLP profile ↗
35ranked-venue papers
1as first author
8since 2021 · last 2024
0009-0002-8856-2472ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
19 papers |
Language models and text generation · 38% Question answering and dialogue systems · 36% Robot navigation and mapping · 9% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 98% Machine learning and data management · 2% |
Topics — the 30 heaviest of 40, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation |
0.7 | 2 | 2019 | Structuring Latent Spaces for Stylized Response Generation · EMNLP/IJCNLP (1) 2019 Generating Informative and Diverse Conversational Responses via Adversarial Information Maximization · NeurIPS 2018 |
Natural language and speech › Question answering and dialogue systems
knowledge-grounded dialogue |
0.7 | 2 | 2019 | Conversing by Reading: Contentful Neural Conversation with On-demand Machine Reading · ACL (1) 2019 A Knowledge-Grounded Neural Conversation Model · AAAI 2018 |
Natural language and speech › Language models and text generation › text generation
grounded text generation |
0.6 | 1 | 2022 | RetGen: A Joint Framework for Retrieval and Grounded Text Generation Modeling · AAAI 2022 |
Natural language and speech › Language models and text generation
hallucination detection |
0.6 | 1 | 2022 | A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text Generation · ACL (1) 2022 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.6 | 1 | 2022 | RetGen: A Joint Framework for Retrieval and Grounded Text Generation Modeling · AAAI 2022 |
Information retrieval
retrieval models |
0.6 | 1 | 2022 | RetGen: A Joint Framework for Retrieval and Grounded Text Generation Modeling · AAAI 2022 |
Natural language and speech › Language models and text generation
controllable text generation |
0.5 | 1 | 2021 | A Controllable Model of Grounded Response Generation · AAAI 2021 |
Natural language and speech › Question answering and dialogue systems › dialogue generation › dialogue response generation
knowledge-grounded response generation |
0.5 | 1 | 2021 | A Controllable Model of Grounded Response Generation · AAAI 2021 |
Natural language and speech › Question answering and dialogue systems
dialogue generation |
0.5 | 2 | 2020 | Emotional Dialogue Generation using Image-Grounded Language Models · CHI 2018 Dialogue Response Ranking Training with Large-Scale Human Feedback Data · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › text generation
constrained text generation |
0.4 | 1 | 2020 | POINTER: Constrained Progressive Text Generation via Insertion-based Generative Pre-training · EMNLP (1) 2020 |
Computer vision › Vision and language
cross-modal alignment |
0.4 | 1 | 2020 | A Recipe for Creating Multimodal Aligned Datasets for Sequential Tasks · ACL 2020 |
Natural language and speech › Language models and text generation › alignment
instruction alignment |
0.4 | 1 | 2020 | A Recipe for Creating Multimodal Aligned Datasets for Sequential Tasks · ACL 2020 |
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
non-autoregressive generation |
0.4 | 1 | 2020 | POINTER: Constrained Progressive Text Generation via Insertion-based Generative Pre-training · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems
response selection |
0.4 | 1 | 2020 | Dialogue Response Ranking Training with Large-Scale Human Feedback Data · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems › dialogue modeling
neural conversation model |
0.4 | 2 | 2018 | A Knowledge-Grounded Neural Conversation Model · AAAI 2018 A Persona-Based Neural Conversation Model · ACL (1) 2016 |
Robotics › Robot navigation and mapping
embodied navigation |
0.4 | 1 | 2019 | Vision-Based Navigation With Language-Based Assistance via Imitation Learning With Indirect Intervention · CVPR 2019 |
Machine learning › Reinforcement learning
imitation learning |
0.4 | 1 | 2019 | Vision-Based Navigation With Language-Based Assistance via Imitation Learning With Indirect Intervention · CVPR 2019 |
Robotics › Robot navigation and mapping › visual navigation
language-guided navigation |
0.4 | 1 | 2019 | Vision-Based Navigation With Language-Based Assistance via Imitation Learning With Indirect Intervention · CVPR 2019 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.4 | 1 | 2019 | Conversing by Reading: Contentful Neural Conversation with On-demand Machine Reading · ACL (1) 2019 |
Natural language and speech › Language models and text generation › controllable text generation
stylized response generation |
0.4 | 1 | 2019 | Structuring Latent Spaces for Stylized Response Generation · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation › controllable text generation
text style transfer |
0.4 | 1 | 2019 | Domain Adaptive Text Style Transfer · EMNLP/IJCNLP (1) 2019 |
Robotics › Robot navigation and mapping
visual navigation |
0.4 | 1 | 2019 | Vision-Based Navigation With Language-Based Assistance via Imitation Learning With Indirect Intervention · CVPR 2019 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.3 | 1 | 2018 | Generating Informative and Diverse Conversational Responses via Adversarial Information Maximization · NeurIPS 2018 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
emotional conversation generation |
0.3 | 1 | 2018 | Emotional Dialogue Generation using Image-Grounded Language Models · CHI 2018 |
Natural language and speech › Question answering and dialogue systems › dialogue generation › dialogue response generation
neural response generation |
0.3 | 1 | 2017 | Steering Output Style and Topic in Neural Response Generation · EMNLP 2017 |
Natural language and speech › Question answering and dialogue systems › personalized dialogue
persona-grounded dialogue |
0.2 | 1 | 2016 | A Persona-Based Neural Conversation Model · ACL (1) 2016 |
Natural language and speech › Language models and text generation
text summarization |
0.2 | 1 | 2016 | A Dataset and Evaluation Metrics for Abstractive Compression of Sentences and Short Paragraphs · EMNLP 2016 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation › grounding
knowledge grounding |
0.1 | 1 | 2021 | A Controllable Model of Grounded Response Generation · AAAI 2021 |
Natural language and speech › Question answering and dialogue systems
conversational agents |
0.1 | 1 | 2017 | Steering Output Style and Topic in Neural Response Generation · EMNLP 2017 |
Natural language and speech › Language models and text generation › text generation
text simplification |
0.1 | 1 | 2016 | A Dataset and Evaluation Metrics for Abstractive Compression of Sentences and Short Paragraphs · EMNLP 2016 |
Methods — techniques the papers use, named apart from their topics
mixture of experts · 1.1joint training · 1.1adversarial training · 0.7neural conversation model · 0.6reference-free detection · 0.6language model pretraining · 0.6language model pre-training · 0.6transformer · 0.5inductive attention · 0.5control phrase prediction · 0.5graph algorithms · 0.4machine learning · 0.0classification · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | GENEVA: GENErating and Visualizing branching narratives using LLMsabstractDialogue-based Role Playing Games (RPGs) require powerful storytelling. The narratives of these may take years to write and typically involve a large creative team. In this work, we demonstrate the potential of large generative text models to assist this process. GENEVA, a prototype tool, generates a rich narrative graph with branching and reconverging storylines that match a high-level narrative description and constraints provided by the designer. A large language model (LLM), GPT-4, is used to generate the branching narrative and to render it in a graph format in a two-step process. We illustrate the use of GENEVA in generating new branching narratives for four well-known stories under different contextual constraints. This tool has the potential to assist in game development, simulations, and other applications with game-like properties. Jorge Leandro, Sudha Rao, Michael Xu, Weijia Xu, Nebojsa Jojic, Chris Brockett, William B. Dolan |
CoG | 6 |
| 2024 | Player-Driven Emergence in LLM-Driven Game NarrativeabstractWe explore how interaction with large language models (LLMs) can give rise to emergent behaviors, empowering players to participate in the evolution of game narratives. Our testbed is a text-adventure game in which players attempt to solve a mystery under a fixed narrative premise, but can freely interact with non-player characters generated by GPT-4, a large language model. We recruit 28 gamers to play the game and use GPT-4 to automatically convert the game logs into a node-graph representing the narrative in the player’s gameplay. We find that through their interactions with the non-deterministic behavior of the LLM, players are able to discover interesting new emergent nodes that were not a part of the original narrative but have potential for being fun and engaging. Players that created the most emergent nodes tended to be those that often enjoy games that facilitate discovery, exploration and experimentation. Jessica Quaye, Sudha Rao, Weijia Xu, Portia Botchway, Chris Brockett, Nebojsa Jojic, Gabriel DesGarennes, Ken Lobb, Michael Xu, Jorge Leandro, Claire Jin, William B. Dolan |
CoG | 6 |
| 2024 | Investigating Agency of LLMs in Human-AI Collaboration TasksabstractAshish Sharma, Sudha Rao, Chris Brockett, Akanksha Malhotra, Nebojsa Jojic, Bill Dolan. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Sudha Rao, Chris Brockett, Akanksha Malhotra, Nebojsa Jojic, William B. Dolan |
EACL (1) | 3 |
| 2022 | RetGen: A Joint Framework for Retrieval and Grounded Text Generation ModelingabstractRecent advances in large-scale pre-training such as GPT-3 allow seemingly high quality text to be generated from a given prompt. However, such generation systems often suffer from problems of hallucinated facts, and are not inherently designed to incorporate useful external information. Grounded generation models appear to offer remedies, but their training typically relies on rarely-available parallel data where information-relevant documents are provided for context. We propose a framework that alleviates this data constraint by jointly training a grounded generator and document retriever on the language model signal. The model learns to reward retrieval of the documents with the highest utility in generation, and attentively combines them using a Mixture-of-Experts (MoE) ensemble to generate follow-on text. We demonstrate that both generator and retriever can take advantage of this joint training and work synergistically to produce more informative and relevant text in both prose and dialogue generation. Yizhe Zhang 0002, Xiang Gao 0011, Yuwei Fang, Chris Brockett, Michel Galley, Jianfeng Gao 0001, William B. Dolan |
AAAI | 5 |
| 2022 | A Token-level Reference-free Hallucination Detection Benchmark for Free-form Text GenerationabstractTianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, Bill Dolan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Tianyu Liu 0001, Yizhe Zhang 0002, Chris Brockett, Zhifang Sui, Weizhu Chen, William B. Dolan |
ACL (1) | 3 |
| 2021 | A Controllable Model of Grounded Response GenerationabstractCurrent end-to-end neural conversation models inherently lack the flexibility to impose semantic control in the response generation process, often resulting in uninteresting responses. Attempts to boost informativeness alone come at the expense of factual accuracy, as attested by pretrained language models' propensity to "hallucinate" facts. While this may be mitigated by access to background knowledge, there is scant guarantee of relevance and informativeness in generated responses. We propose a framework that we call controllable grounded response generation (CGRG), in which lexical control phrases are either provided by a user or automatically extracted by a control phrase predictor from dialogue context and grounding knowledge. Quantitative and qualitative results show that, using this framework, a transformer based model with a novel inductive attention mechanism, trained on a conversation-like Reddit dataset, outperforms strong generation baselines. Zeqiu Wu, Michel Galley, Chris Brockett, Yizhe Zhang 0002, Xiang Gao 0011, Chris Quirk, Rik Koncel-Kedziorski, Jianfeng Gao 0001, Hannaneh Hajishirzi, Mari Ostendorf, William B. Dolan |
AAAI | 3 |
| 2021 | Text Editing by CommandabstractFelix Faltings, Michel Galley, Gerold Hintz, Chris Brockett, Chris Quirk, Jianfeng Gao, Bill Dolan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Felix Faltings, Michel Galley, Gerold Hintz, Chris Brockett, Chris Quirk, Jianfeng Gao 0001, William B. Dolan |
NAACL-HLT | 4 |
| 2021 | Contextualized Perturbation for Textual Adversarial AttackabstractDianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, Bill Dolan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Dianqi Li, Yizhe Zhang 0002, Hao Peng 0009, Liqun Chen 0001, Chris Brockett, Ming-Ting Sun, William B. Dolan |
NAACL-HLT | 5 |
| 2020 | A Recipe for Creating Multimodal Aligned Datasets for Sequential TasksabstractMany high-level procedural tasks can be decomposed into sequences of instructions that vary in their order and choice of tools. In the cooking domain, the web offers many partially-overlapping text and video recipes (i.e. procedures) that describe how to make the same dish (i.e. high-level task). Aligning instructions for the same dish across different sources can yield descriptive visual explanations that are far richer semantically than conventional textual instructions, providing commonsense insight into how real-world procedures are structured. Learning to align these different instruction sets is challenging because: a) different recipes vary in their order of instructions and use of ingredients; and b) video instructions can be noisy and tend to contain far more information than text instructions. To address these challenges, we first use an unsupervised alignment algorithm that learns pairwise alignments between instructions of different recipes for the same dish. We then use a graph algorithm to derive a joint alignment between multiple text and multiple video recipes for the same dish. We release the Microsoft Research Multimodal Aligned Recipe Corpus containing 150K pairwise alignments between recipes across 4,262 dishes with rich commonsense information. Angela S. Lin, Sudha Rao, Asli Celikyilmaz, Elnaz Nouri, Chris Brockett, Debadeepta Dey, William B. Dolan |
ACL | 5 |
| 2020 | Dialogue Response Ranking Training with Large-Scale Human Feedback DataabstractExisting open-domain dialog models are generally trained to minimize the perplexity of target human responses.However, some human replies are more engaging than others, spawning more followup interactions.Current conversational models are increasingly capable of producing turns that are context-relevant, but in order to produce compelling agents, these models need to be able to predict and optimize for turns that are genuinely engaging.We leverage social media feedback data (number of replies and upvotes) to build a large-scale training dataset for feedback prediction.To alleviate possible distortion between the feedback and engagingness, we convert the ranking problem to a comparison of response pairs which involve few confounding factors.We trained DIALOGRPT, a set of GPT-2 based models on 133M pairs of human feedback data and the resulting ranker outperformed several baselines.Particularly, our ranker outperforms the conventional dialog perplexity baseline with a large margin on predicting Reddit feedback.We finally combine the feedback prediction models and a human-like scoring model to rank the machine-generated dialog responses.Crowd-sourced human evaluation shows that our ranking method correlates better with real human preferences than baseline models. 1 Xiang Gao 0011, Yizhe Zhang 0002, Michel Galley, Chris Brockett, William B. Dolan |
EMNLP (1) | 4 |
| 2020 | POINTER: Constrained Progressive Text Generation via Insertion-based Generative Pre-trainingabstractLarge-scale pre-trained language models, such as BERT and GPT-2, have achieved excellent performance in language representation learning and free-form text generation.However, these models cannot be directly employed to generate text under specified lexical constraints.To address this challenge, we present POINTER 1 , a simple yet novel insertion-based approach for hard-constrained text generation.The proposed method operates by progressively inserting new tokens between existing tokens in a parallel manner.This procedure is recursively applied until a sequence is completed.The resulting coarse-to-fine hierarchy makes the generation process intuitive and interpretable.We pre-train our model with the proposed progressive insertion-based objective on a 12GB Wikipedia dataset, and finetune it on downstream hard-constrained generation tasks.Non-autoregressive decoding yields an empirically logarithmic time complexity during inference time.Experimental results on both News and Yelp datasets demonstrate that POINTER achieves state-of-the-art performance on constrained text generation.We released the pre-trained models and the source code to facilitate future research 2 . Yizhe Zhang 0002, Guoyin Wang 0002, Chunyuan Li, Zhe Gan, Chris Brockett, William B. Dolan |
EMNLP (1) | 5 |
| 2019 | Conversing by Reading: Contentful Neural Conversation with On-demand Machine ReadingabstractLianhui Qin, Michel Galley, Chris Brockett, Xiaodong Liu, Xiang Gao, Bill Dolan, Yejin Choi, Jianfeng Gao. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Lianhui Qin, Michel Galley, Chris Brockett, Xiaodong Liu 0003, Xiang Gao 0011, William B. Dolan, Yejin Choi 0001, Jianfeng Gao 0001 |
ACL (1) | 3 |
| 2019 | Vision-Based Navigation With Language-Based Assistance via Imitation Learning With Indirect InterventionabstractWe present Vision-based Navigation with Language-based Assistance (VNLA), a grounded vision-language task where an agent with visual perception is guided via language to find objects in photorealistic indoor environments. The task emulates a real-world scenario in that (a) the requester may not know how to navigate to the target objects and thus makes requests by only specifying high-level end-goals, and (b) the agent is capable of sensing when it is lost and querying an advisor, who is more qualified at the task, to obtain language subgoals to make progress. To model language-based assistance, we develop a general framework termed Imitation Learning with Indirect Intervention (I3L), and propose a solution that is effective on the VNLA task. Empirical results show that this approach significantly improves the success rate of the learning agent over other baselines on both seen and unseen environments. Our code and data are publicly available at https://github.com/debadeepta/vnla . Debadeepta Dey, Chris Brockett, William B. Dolan |
CVPR | 3 |
| 2019 | Structuring Latent Spaces for Stylized Response GenerationabstractXiang Gao, Yizhe Zhang, Sungjin Lee, Michel Galley, Chris Brockett, Jianfeng Gao, Bill Dolan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiang Gao 0011, Yizhe Zhang 0002, Michel Galley, Chris Brockett, Jianfeng Gao 0001, William B. Dolan |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Domain Adaptive Text Style TransferabstractDianqi Li, Yizhe Zhang, Zhe Gan, Yu Cheng, Chris Brockett, Bill Dolan, Ming-Ting Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dianqi Li, Yizhe Zhang 0002, Zhe Gan, Yu Cheng 0001, Chris Brockett, William B. Dolan, Ming-Ting Sun |
EMNLP/IJCNLP (1) | 5 |
| 2018 | A Knowledge-Grounded Neural Conversation ModelabstractNeural network models are capable of generating extremely natural sounding conversational interactions. However, these models have been mostly applied to casual scenarios (e.g., as “chatbots”) and have yet to demonstrate they can serve in more useful conversational applications. This paper presents a novel, fully data-driven, and knowledge-grounded neural conversation model aimed at producing more contentful responses. We generalize the widely-used Sequence-to-Sequence (Seq2Seq) approach by conditioning responses on both conversation history and external “facts”, allowing the model to be versatile and applicable in an open-domain setting. Our approach yields significant improvements over a competitive Seq2Seq baseline. Human judges found that our outputs are significantly more informative. Marjan Ghazvininejad, Chris Brockett, Ming-Wei Chang, William B. Dolan, Jianfeng Gao 0001, Scott Yih, Michel Galley |
AAAI | 2 |
| 2018 | Emotional Dialogue Generation using Image-Grounded Language ModelsabstractComputer-based conversational agents are becoming ubiquitous. However, for these systems to be engaging and valuable to the user, they must be able to express emotion, in addition to providing informative responses. Humans rely on much more than language during conversations; visual information is key to providing context. We present the first example of an image-grounded conversational agent using visual sentiment, facial expression and scene features. We show that key qualities of the generated dialogue can be manipulated by the features used for training the agent. We evaluate our model on a large and very challenging real-world dataset of conversations from social media (Twitter). The image-grounding leads to significantly more informative, emotional and specific responses, and the exact qualities can be tuned depending on the image features used. Furthermore, our model improves the objective quality of dialogue responses when evaluated on standard natural language metrics. Bernd Huber, Daniel McDuff, Chris Brockett, Michel Galley, William B. Dolan |
CHI | 3 |
| 2018 | Generating Informative and Diverse Conversational Responses via Adversarial Information MaximizationabstractResponses generated by neural conversational models tend to lack informativeness and diversity. We present Adversarial Information Maximization (AIM), an adversarial learning framework that addresses these two related but distinct problems. To foster response diversity, we leverage adversarial training that allows distributional matching of synthetic and real responses. To improve informativeness, our framework explicitly optimizes a variational lower bound on pairwise mutual information between query and response. Empirical results from automatic and human evaluations demonstrate that our methods significantly boost informativeness and diversity. Yizhe Zhang 0002, Michel Galley, Jianfeng Gao 0001, Zhe Gan, Xiujun Li, Chris Brockett, William B. Dolan |
NeurIPS | 6 |
| 2017 | Steering Output Style and Topic in Neural Response GenerationabstractWe propose simple and flexible training and decoding methods for influencing output style and topic in neural encoderdecoder based language generation.This capability is desirable in a variety of applications, including conversational systems, where successful agents need to produce language in a specific style and generate responses steered by a human puppeteer or external knowledge.We decompose the neural generation process into empirically easier sub-problems: a faithfulness model and a decoding method based on selectivesampling.We also describe training and sampling algorithms that bias the generation process with a specific language style restriction, or a topic restriction.Human evaluation results show that our proposed methods are able to to restrict style and topic without degrading output quality in conversational tasks.chat context Where are you?Twitter bot i'm on my way to london now.Star Wars bot we're heading to the dark side of jabba's palace.Hillary bot i'm here in philadelphia, pennsylvania Trump bot i went to the white house.we were in brooklyn.Kennedy bot i am on the question of the united states.chat context What are you doing?Twitter bot i'm just chilling with my dad, i'm just sitting in the kitchen and listening to music Star Wars bot i'm doing my magic solo.Hillary bot i'm running for president, i'm going to be talking about some of these things Trump bot i'm building a wall, right now Kennedy bot today, i am asking the congress for a new program to make a new effort to increase the tax privileges and to stimulate Di Wang 0030, Nebojsa Jojic, Chris Brockett, Eric Nyberg |
EMNLP | 3 |
| 2017 | Multi-Task Learning for Speaker-Role Adaptation in Neural Conversation ModelsabstractBuilding a persona-based conversation agent is challenging owing to the lack of large amounts of speaker-specific conversation data for model training. This paper addresses the problem by proposing a multi-task learning approach to training neural conversation models that leverages both conversation data across speakers and other types of data pertaining to the speaker and speaker roles to be modeled. Experiments show that our approach leads to significant improvements over baseline model quality, generating responses that capture more precisely speakers’ traits and speaking styles. The model offers the benefits of being algorithmically simple and easy to implement, and not relying on large quantities of data representing specific individual speakers. Yi Luan, Chris Brockett, William B. Dolan, Jianfeng Gao 0001, Michel Galley |
IJCNLP(1) | 2 |
| 2017 | Image-Grounded Conversations: Multimodal Context for Natural Question and Response GenerationabstractThe popularity of image sharing on social media and the engagement it creates between users reflect the important role that visual context plays in everyday conversations. We present a novel task, Image Grounded Conversations (IGC), in which natural-sounding conversations are generated about a shared image. To benchmark progress, we introduce a new multiple reference dataset of crowd-sourced, event-centric conversations on images. IGC falls on the continuum between chit-chat and goal-directed conversation models, where visual grounding constrains the topic of conversation to event-driven utterances. Experiments with models trained on social media data show that the combination of visual and textual context enhances the quality of generated conversational turns. In human evaluation, the gap between human performance and that of both neural and retrieval architectures suggests that multi-modal IGC presents an interesting challenge for dialog research. Nasrin Mostafazadeh, Chris Brockett, William B. Dolan, Michel Galley, Jianfeng Gao 0001, Georgios Spithourakis, Lucy Vanderwende |
IJCNLP(1) | 2 |
| 2016 | A Persona-Based Neural Conversation ModelabstractJiwei Li, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao, Bill Dolan. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Jiwei Li 0001, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao 0001, William B. Dolan |
ACL (1) | 3 |
| 2016 | A Dataset and Evaluation Metrics for Abstractive Compression of Sentences and Short ParagraphsabstractWe introduce a manually-created, multireference dataset for abstractive sentence and short paragraph compression.First, we examine the impact of single-and multi-sentence level editing operations on human compression quality as found in this corpus.We observe that substitution and rephrasing operations are more meaning preserving than other operations, and that compressing in context improves quality.Second, we systematically explore the correlations between automatic evaluation metrics and human judgments of meaning preservation and grammaticality in the compression task, and analyze the impact of the linguistic units used and precision versus recall measures on the quality of the metrics.Multi-reference evaluation metrics are shown to offer significant advantage over single reference-based metrics. Kristina Toutanova, Chris Brockett, Ke M. Tran, Saleema Amershi |
EMNLP | 2 |
| 2016 | A Diversity-Promoting Objective Function for Neural Conversation ModelsabstractJiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, Bill Dolan. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Jiwei Li 0001, Michel Galley, Chris Brockett, Jianfeng Gao 0001, William B. Dolan |
HLT-NAACL | 3 |
| 2015 | A Neural Network Approach to Context-Sensitive Generation of Conversational ResponsesabstractAlessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, Bill Dolan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao 0001, William B. Dolan |
HLT-NAACL | 4 |
| 2013 | Paraphrase features to improve natural language understanding
Ruhi Sarikaya, Chris Brockett, Chris Quirk, William B. Dolan |
INTERSPEECH | 3 |
| 2010 | Hitting the Right Paraphrases in Good Time
Stanley Kok, Chris Brockett |
HLT-NAACL | 2 |
| 2008 | Using Contextual Speller Techniques and Language Modeling for ESL Error Correction
Michael Gamon, Jianfeng Gao 0001, Chris Brockett, Alexandre Klementiev, William B. Dolan, Dmitriy Belenko, Lucy Vanderwende |
IJCNLP | 3 |
| 2007 | Beyond SumBasic: Task-focused summarization with sentence simplification and lexical expansion
Lucy Vanderwende, Hisami Suzuki, Chris Brockett, Ani Nenkova |
Inf. Process. Manag. | 3 |
| 2006 | Correcting ESL Errors Using Phrasal SMT TechniquesabstractThis paper presents a pilot study of the use of phrasal Statistical Machine Translation (SMT) techniques to identify and correct writing errors made by learners of English as a Second Language (ESL). Using examples of mass noun errors found in the Chinese Learner Error Corpus (CLEC) to guide creation of an engineered training set, we show that application of the SMT paradigm can capture errors not well addressed by widely-used proofing tools designed for native speakers. Our system was able to correct 61.81% of mistakes in a set of naturally-occurring examples of mass noun errors found on the World Wide Web, suggesting that efforts to collect alignable corpora of pre- and post-editing ESL writing samples offer can enable the development of SMT-based writing assistance tools capable of repairing many of the complex syntactic and lexical problems found in the writing of ESL learners. Chris Brockett, William B. Dolan, Michael Gamon |
ACL | 1 |
| 2004 | Unsupervised Construction of Large Paraphrase Corpora: Exploiting Massively Parallel News Sources
William B. Dolan, Chris Quirk, Chris Brockett |
COLING | 3 |
| 2004 | Monolingual Machine Translation for Paraphrase Generation
Chris Quirk, Chris Brockett, William B. Dolan |
EMNLP | 2 |
| 2001 | A Machine Learning Approach to the Automatic Evaluation of Machine TranslationabstractWe present a machine learning approach to evaluating the well-formedness of output of a machine translation system, using classifiers that learn to distinguish human reference translations from machine translations. This approach can be used to evaluate an MT system, tracking improvements over time; to aid in the kind of failure analysis that can help guide system development; and to select among alternative output strings. The method presented is fully automated and independent of source language, target language and domain. Simon Corston-Oliver, Michael Gamon, Chris Brockett |
ACL | 3 |
| 2000 | Robust Segmentation of Japanese Text into a Lattice for Parsing
Gary Kacmarcik, Chris Brockett, Hisami Suzuki |
COLING | 2 |
| 2000 | Using a Broad-Coverage Parser for Word-Breaking in Japanese
Hisami Suzuki, Chris Brockett, Gary Kacmarcik |
COLING | 2 |