EDBT 2026 Demo / reviewers in the wild / expert
Michel Galley
dblp:05/3289
· DBLP profile ↗
56ranked-venue papers
14as first author
17since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 52 · 14 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?abstractYao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu, Jianfeng Gao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu 0004, Jianfeng Gao 0001 |
EMNLP | 2 |
| 2025 | ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory LearningabstractAutonomous agents have demonstrated significant potential in automating complex multistep decision-making tasks. However, even state-of-the-art vision-language models (VLMs), such as GPT-4o, still fall short of human-level performance, particularly in intricate web environments and long-horizon planning tasks. To address these limitations, we introduce Reflective Monte Carlo Tree Search (R-MCTS), a novel test-time algorithm designed to enhance the ability of AI agents, e.g., powered by GPT-4o, to explore decision space on the fly.
R-MCTS extends traditional MCTS by 1) incorporating contrastive reflection, allowing agents to learn from past interactions and dynamically improve their search efficiency; and 2) using multi-agent debate to provide reliable state evaluation. Moreover, we improve the agent's performance by fine-tuning GPT-4o through self-learning, using R-MCTS generated tree traversals without any human-provided labels. On the challenging VisualWebArena benchmark, our GPT-4o-based R-MCTS agent achieves a 6% to 30% relative improvement across various tasks compared to the previous state-of-the-art. Additionally, we show that the knowledge gained from test-time search can be effectively transferred back to GPT-4o via fine-tuning. The fine-tuned GPT-4o matches 97\% of R-MCTS's performance while reducing compute usage by a factor of four at test time. Furthermore, qualitative results reveal that the fine-tuned GPT-4o model demonstrates the ability to explore the environment, evaluate a state, and backtrack to viable ones when it detects that the current state cannot lead to success. Moreover, our work demonstrates the compute scaling properties in both training - data collection with R-MCTS - and testing time. These results suggest a promising research direction to enhance VLMs' reasoning and planning capabilities for agentic applications via test-time search and self-learning. Xiao Yu 0011, Baolin Peng, Vineeth Vajipey, Hao Cheng 0002, Michel Galley, Jianfeng Gao 0001, Zhou Yu 0005 |
ICLR | 5 |
| 2025 | CollabLLM: From Passive Responders to Active CollaboratorsabstractLarge Language Models are typically trained with next-turn rewards, limiting their ability to optimize for long-term interaction. As a result, they often respond passively to ambiguous or open-ended user requests, failing to help users reach their ultimate intents and leading to inefficient conversations. To address these limitations, we introduce CollabLLM, a novel and general training framework that enhances multiturn human-LLM collaboration. Its key innovation is a collaborative simulation that estimates the long-term contribution of responses
using Multiturn-aware Rewards. By reinforcement fine-tuning these rewards, CollabLLM goes beyond responding to user requests, and actively uncovers user intent and offers insightful suggestions—a key step towards more human-centered AI. We also devise a multiturn interaction benchmark with three challenging tasks such as document creation. CollabLLM significantly outperforms our baselines with averages of 18.5% higher task performance and 46.3% improved interactivity by LLM judges. Finally, we conduct a large user study with 201 judges, where CollabLLM increases user satisfaction by 17.6% and reduces user spent time by 10.4%. Shirley Wu, Michel Galley, Baolin Peng, Hao Cheng 0002, Gavin Li, Yao Dou, Weixin Cai, James Zou 0001, Jure Leskovec, Jianfeng Gao 0001 |
ICML | 2 |
| 2025 | Iterative Self-Tuning LLMs for Enhanced Jailbreaking CapabilitiesabstractChung-En Sun, Xiaodong Liu, Weiwei Yang, Tsui-Wei Weng, Hao Cheng, Aidan San, Michel Galley, Jianfeng Gao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Chung-En Sun, Xiaodong Liu 0003, Tsui-Wei Weng, Hao Cheng 0002, Aidan San, Michel Galley, Jianfeng Gao 0001 |
NAACL (Long Papers) | 7 |
| 2024 | MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsabstractLarge Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understanding and compositional reasoning, which all state-of-the-art foundation models find challenging. With MathVista, we have conducted a comprehensive, quantitative evaluation of 12 prominent foundation models. The best-performing GPT-4V model achieves an overall accuracy of 49.9%, substantially outperforming Bard, the second-best performer, by 15.1%. Our in-depth analysis reveals that the superiority of GPT-4V is mainly attributed to its enhanced visual perception and mathematical reasoning. However, GPT-4V still falls short of human performance by 10.4%, as it often struggles to understand complex figures and perform rigorous reasoning. This significant gap underscores the critical role that MathVista will play in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. We further explore the new ability of self-verification, the application of self-consistency, and the interactive chatbot capabilities of GPT-4V, highlighting its promising potential for future research. The project is available at https://mathvista.github.io/. Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 0010, Chunyuan Li, Hannaneh Hajishirzi, Hao Cheng 0002, Kai-Wei Chang 0001, Michel Galley, Jianfeng Gao 0001 |
ICLR | 9 |
| 2024 | Teaching Language Models to Self-Improve through Interactive DemonstrationsabstractXiao Yu, Baolin Peng, Michel Galley, Jianfeng Gao, Zhou Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xiao Yu 0011, Baolin Peng, Michel Galley, Jianfeng Gao 0001, Zhou Yu 0005 |
NAACL-HLT | 3 |
| 2023 | DIONYSUS: A Pre-trained Model for Low-Resource Dialogue SummarizationabstractDialogue summarization has recently garnered significant attention due to its wide range of applications.However, existing methods for summarizing dialogues have limitations because they do not take into account the inherent structure of dialogue and rely heavily on labeled data, which can lead to poor performance in new domains.In this work, we propose DIONYSUS (dynamic input optimization in pre-training for dialogue summarization), a pre-trained encoder-decoder model for summarizing dialogues in any new domain.To pretrain DIONYSUS, we create two pseudo summaries for each dialogue example: one from a fine-tuned summarization model and the other from important dialogue turns.We then choose one of these pseudo summaries based on information distribution differences in different types of dialogues.This selected pseudo summary serves as the objective for pre-training DIONYSUS using a self-supervised approach on a large dialogue corpus.Our experiments show that DIONYSUS outperforms existing methods on six datasets, as demonstrated by its ROUGE scores in zero-shot and few-shot settings. Yu Li 0013, Baolin Peng, Michel Galley, Zhou Yu 0005, Jianfeng Gao 0001 |
ACL (1) | 4 |
| 2023 | Interactive Text GenerationabstractFelix Faltings, Michel Galley, Kianté Brantley, Baolin Peng, Weixin Cai, Yizhe Zhang, Jianfeng Gao, Bill Dolan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Felix Faltings, Michel Galley, Kianté Brantley, Baolin Peng, Weixin Cai, Yizhe Zhang 0002, Jianfeng Gao 0001, William B. Dolan |
EMNLP | 2 |
| 2023 | Guiding Large Language Models via Directional Stimulus PromptingabstractWe introduce Directional Stimulus Prompting, a novel framework for guiding black-box large language models (LLMs) towards specific desired outputs. Instead of directly adjusting LLMs, our method employs a small tunable policy model (e.g., T5) to generate an auxiliary directional stimulus prompt for each input instance. These directional stimulus prompts act as nuanced, instance-specific hints and clues to guide LLMs in generating desired outcomes, such as including specific keywords in the generated summary. Our approach sidesteps the challenges of direct LLM tuning by optimizing the policy model to explore directional stimulus prompts that align LLMs with desired behaviors. The policy model can be optimized through 1) supervised fine-tuning using labeled data and 2) reinforcement learning from offline or online rewards based on the LLM's output. We evaluate our method across various tasks, including summarization, dialogue response generation, and chain-of-thought reasoning. Our experiments indicate a consistent improvement in the performance of LLMs such as ChatGPT, Codex, and InstructGPT on these supervised tasks with minimal labeled data. Remarkably, by utilizing merely 80 dialogues from the MultiWOZ dataset, our approach boosts ChatGPT's performance by a relative 41.4%, achieving or exceeding the performance of some fully supervised state-of-the-art models. Moreover, the instance-specific chain-of-thought prompt generated through our method enhances InstructGPT's reasoning accuracy, outperforming both generalized human-crafted prompts and those generated through automatic prompt engineering. The code and data are publicly available at https://github.com/Leezekun/Directional-Stimulus-Prompting. Zekun Li 0001, Baolin Peng, Michel Galley, Jianfeng Gao 0001, Xifeng Yan |
NeurIPS | 4 |
| 2023 | Chameleon: Plug-and-Play Compositional Reasoning with Large Language ModelsabstractLarge language models (LLMs) have achieved remarkable progress in solving various natural language processing tasks due to emergent reasoning abilities. However, LLMs have inherent limitations as they are incapable of accessing up-to-date information (stored on the Web or in task-specific knowledge bases), using external tools, and performing precise mathematical and logical reasoning. In this paper, we present Chameleon, an AI system that mitigates these limitations by augmenting LLMs with plug-and-play modules for compositional reasoning. Chameleon synthesizes programs by composing various tools (e.g., LLMs, off-the-shelf vision models, web search engines, Python functions, and heuristic-based modules) for accomplishing complex reasoning tasks. At the heart of Chameleon is an LLM-based planner that assembles a sequence of tools to execute to generate the final response. We showcase the effectiveness of Chameleon on two multi-modal knowledge-intensive reasoning tasks: ScienceQA and TabMWP. Chameleon, powered by GPT-4, achieves an 86.54% overall accuracy on ScienceQA, improving the best published few-shot result by 11.37%. On TabMWP, GPT-4-powered Chameleon improves the accuracy by 17.0%, lifting the state of the art to 98.78%. Our analysis also shows that the GPT-4-powered planner exhibits more consistent and rational tool selection via inferring potential constraints from instructions, compared to a ChatGPT-powered planner. Pan Lu, Baolin Peng, Hao Cheng 0002, Michel Galley, Kai-Wei Chang 0001, Ying Nian Wu, Song-Chun Zhu, Jianfeng Gao 0001 |
NeurIPS | 4 |
| 2023 | Enhancing Task Bot Engagement with Synthesized Open-Domain DialogabstractThe construction of dialog systems for various types of conversations, such as task-oriented dialog (TOD) and open-domain dialog (ODD), has been an active area of research.In order to more closely mimic human-like conversations that often involve the fusion of different dialog modes, it is important to develop systems that can effectively handle both TOD and ODD and access different knowledge sources.In this work, we present a new automatic framework to enrich TODs with synthesized ODDs.We also introduce the PivotBot model, which is capable of handling both TOD and ODD modes and can access different knowledge sources to generate informative responses.Evaluation results indicate the superior ability of the proposed model to switch smoothly between TOD and ODD tasks. Miaoran Li, Baolin Peng, Michel Galley, Jianfeng Gao 0001, Zhu (Drew) Zhang |
SIGDIAL | 3 |
| 2022 | RetGen: A Joint Framework for Retrieval and Grounded Text Generation ModelingabstractRecent advances in large-scale pre-training such as GPT-3 allow seemingly high quality text to be generated from a given prompt. However, such generation systems often suffer from problems of hallucinated facts, and are not inherently designed to incorporate useful external information. Grounded generation models appear to offer remedies, but their training typically relies on rarely-available parallel data where information-relevant documents are provided for context. We propose a framework that alleviates this data constraint by jointly training a grounded generator and document retriever on the language model signal. The model learns to reward retrieval of the documents with the highest utility in generation, and attentively combines them using a Mixture-of-Experts (MoE) ensemble to generate follow-on text. We demonstrate that both generator and retriever can take advantage of this joint training and work synergistically to produce more informative and relevant text in both prose and dialogue generation. Yizhe Zhang 0002, Xiang Gao 0011, Yuwei Fang, Chris Brockett, Michel Galley, Jianfeng Gao 0001, William B. Dolan |
AAAI | 6 |
| 2021 | Data Augmentation for Abstractive Query-Focused Multi-Document SummarizationabstractThe progress in Query-focused Multi-Document Summarization (QMDS) has been limited by the lack of sufficient largescale high-quality training datasets. We present two QMDS training datasets, which we construct using two data augmentation methods: (1) transferring the commonly used single-document CNN/Daily Mail summarization dataset to create the QMDSCNN dataset, and (2) mining search-query logs to create the QMDSIR dataset. These two datasets have complementary properties, i.e., QMDSCNN has real summaries but queries are simulated, while QMDSIR has real queries but simulated summaries. To cover both these real summary and query aspects, we build abstractive end-to-end neural network models on the combined datasets that yield new state-of-the-art transfer results on DUC datasets. We also introduce new hierarchical encoders that enable a more efficient encoding of the query together with multiple documents. Empirical results demonstrate that our data augmentation and encoding methods outperform baseline models on automatic metrics, as well as on human evaluations along multiple attributes. Ramakanth Pasunuru, Asli Celikyilmaz, Michel Galley, Chenyan Xiong, Yizhe Zhang 0002, Mohit Bansal, Jianfeng Gao 0001 |
AAAI | 3 |
| 2021 | A Controllable Model of Grounded Response GenerationabstractCurrent end-to-end neural conversation models inherently lack the flexibility to impose semantic control in the response generation process, often resulting in uninteresting responses. Attempts to boost informativeness alone come at the expense of factual accuracy, as attested by pretrained language models' propensity to "hallucinate" facts. While this may be mitigated by access to background knowledge, there is scant guarantee of relevance and informativeness in generated responses. We propose a framework that we call controllable grounded response generation (CGRG), in which lexical control phrases are either provided by a user or automatically extracted by a control phrase predictor from dialogue context and grounding knowledge. Quantitative and qualitative results show that, using this framework, a transformer based model with a novel inductive attention mechanism, trained on a conversation-like Reddit dataset, outperforms strong generation baselines. Zeqiu Wu, Michel Galley, Chris Brockett, Yizhe Zhang 0002, Xiang Gao 0011, Chris Quirk, Rik Koncel-Kedziorski, Jianfeng Gao 0001, Hannaneh Hajishirzi, Mari Ostendorf, William B. Dolan |
AAAI | 2 |
| 2021 | Text Editing by CommandabstractFelix Faltings, Michel Galley, Gerold Hintz, Chris Brockett, Chris Quirk, Jianfeng Gao, Bill Dolan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Felix Faltings, Michel Galley, Gerold Hintz, Chris Brockett, Chris Quirk, Jianfeng Gao 0001, William B. Dolan |
NAACL-HLT | 2 |
| 2021 | Ask what's missing and what's useful: Improving Clarification Question Generation using Global KnowledgeabstractBodhisattwa Prasad Majumder, Sudha Rao, Michel Galley, Julian McAuley. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Bodhisattwa Prasad Majumder, Sudha Rao, Michel Galley, Julian J. McAuley |
NAACL-HLT | 3 |
| 2021 | Overview of the Eighth Dialog System Technology Challenge: DSTC8abstractThis paper introduces the Eighth Dialog System Technology Challenge. In line with recent challenges, the eighth edition focuses on applying end-to-end dialog technologies in a pragmatic way for multi-domain task-completion, noetic response selection, audio visual scene-aware dialog, and schema-guided dialog state tracking tasks. This paper describes the task definition, provided datasets, baselines and evaluation set-up for each track. We also summarize the results of the submitted systems to highlight the overall trends of the state-of-the-art technologies for the tasks. Seokhwan Kim, Michel Galley, R. Chulaka Gunasekara, Adam Atkinson, Baolin Peng, Hannes Schulz, Jianfeng Gao 0001, Jinchao Li, Mahmoud Adada, Minlie Huang, Luis A. Lastras, Jonathan K. Kummerfeld, Walter S. Lasecki, Chiori Hori, Anoop Cherian, Tim K. Marks, Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Dialogue Response Ranking Training with Large-Scale Human Feedback DataabstractExisting open-domain dialog models are generally trained to minimize the perplexity of target human responses.However, some human replies are more engaging than others, spawning more followup interactions.Current conversational models are increasingly capable of producing turns that are context-relevant, but in order to produce compelling agents, these models need to be able to predict and optimize for turns that are genuinely engaging.We leverage social media feedback data (number of replies and upvotes) to build a large-scale training dataset for feedback prediction.To alleviate possible distortion between the feedback and engagingness, we convert the ranking problem to a comparison of response pairs which involve few confounding factors.We trained DIALOGRPT, a set of GPT-2 based models on 133M pairs of human feedback data and the resulting ranker outperformed several baselines.Particularly, our ranker outperforms the conventional dialog perplexity baseline with a large margin on predicting Reddit feedback.We finally combine the feedback prediction models and a human-like scoring model to rank the machine-generated dialog responses.Crowd-sourced human evaluation shows that our ranking method correlates better with real human preferences than baseline models. 1 Xiang Gao 0011, Yizhe Zhang 0002, Michel Galley, Chris Brockett, William B. Dolan |
EMNLP (1) | 3 |
| 2020 | Overview of the seventh Dialog System Technology Challenge: DSTC7
Luis Fernando D'Haro, Koichiro Yoshino, Chiori Hori, Tim K. Marks, Lazaros Polymenakos, Jonathan K. Kummerfeld, Michel Galley, Xiang Gao 0011 |
Comput. Speech Lang. | 7 |
| 2019 | Conversing by Reading: Contentful Neural Conversation with On-demand Machine ReadingabstractLianhui Qin, Michel Galley, Chris Brockett, Xiaodong Liu, Xiang Gao, Bill Dolan, Yejin Choi, Jianfeng Gao. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Lianhui Qin, Michel Galley, Chris Brockett, Xiaodong Liu 0003, Xiang Gao 0011, William B. Dolan, Yejin Choi 0001, Jianfeng Gao 0001 |
ACL (1) | 2 |
| 2019 | Structuring Latent Spaces for Stylized Response GenerationabstractXiang Gao, Yizhe Zhang, Sungjin Lee, Michel Galley, Chris Brockett, Jianfeng Gao, Bill Dolan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xiang Gao 0011, Yizhe Zhang 0002, Michel Galley, Chris Brockett, Jianfeng Gao 0001, William B. Dolan |
EMNLP/IJCNLP (1) | 4 |
| 2018 | A Knowledge-Grounded Neural Conversation ModelabstractNeural network models are capable of generating extremely natural sounding conversational interactions. However, these models have been mostly applied to casual scenarios (e.g., as “chatbots”) and have yet to demonstrate they can serve in more useful conversational applications. This paper presents a novel, fully data-driven, and knowledge-grounded neural conversation model aimed at producing more contentful responses. We generalize the widely-used Sequence-to-Sequence (Seq2Seq) approach by conditioning responses on both conversation history and external “facts”, allowing the model to be versatile and applicable in an open-domain setting. Our approach yields significant improvements over a competitive Seq2Seq baseline. Human judges found that our outputs are significantly more informative. Marjan Ghazvininejad, Chris Brockett, Ming-Wei Chang, William B. Dolan, Jianfeng Gao 0001, Scott Yih, Michel Galley |
AAAI | 7 |
| 2018 | Emotional Dialogue Generation using Image-Grounded Language ModelsabstractComputer-based conversational agents are becoming ubiquitous. However, for these systems to be engaging and valuable to the user, they must be able to express emotion, in addition to providing informative responses. Humans rely on much more than language during conversations; visual information is key to providing context. We present the first example of an image-grounded conversational agent using visual sentiment, facial expression and scene features. We show that key qualities of the generated dialogue can be manipulated by the features used for training the agent. We evaluate our model on a large and very challenging real-world dataset of conversations from social media (Twitter). The image-grounding leads to significantly more informative, emotional and specific responses, and the exact qualities can be tuned depending on the image features used. Furthermore, our model improves the objective quality of dialogue responses when evaluated on standard natural language metrics. Bernd Huber, Daniel McDuff, Chris Brockett, Michel Galley, William B. Dolan |
CHI | 4 |
| 2018 | Generating Informative and Diverse Conversational Responses via Adversarial Information MaximizationabstractResponses generated by neural conversational models tend to lack informativeness and diversity. We present Adversarial Information Maximization (AIM), an adversarial learning framework that addresses these two related but distinct problems. To foster response diversity, we leverage adversarial training that allows distributional matching of synthetic and real responses. To improve informativeness, our framework explicitly optimizes a variational lower bound on pairwise mutual information between query and response. Empirical results from automatic and human evaluations demonstrate that our methods significantly boost informativeness and diversity. Yizhe Zhang 0002, Michel Galley, Jianfeng Gao 0001, Zhe Gan, Xiujun Li, Chris Brockett, William B. Dolan |
NeurIPS | 2 |
| 2018 | Neural Approaches to Conversational AIabstractThis tutorial surveys neural approaches to conversational AI that were developed in the last few years. We group conversational systems into three categories: (1) question answering agents, (2) task-oriented dialogue agents, and (3) social bots. For each category, we present a review of state-of-the-art neural approaches, draw the connection between neural approaches and traditional symbolic approaches, and discuss the progress we have made and challenges we are facing, using specific systems and models as case studies. Jianfeng Gao 0001, Michel Galley, Lihong Li 0001 |
SIGIR | 2 |
| 2017 | Multi-Task Learning for Speaker-Role Adaptation in Neural Conversation ModelsabstractBuilding a persona-based conversation agent is challenging owing to the lack of large amounts of speaker-specific conversation data for model training. This paper addresses the problem by proposing a multi-task learning approach to training neural conversation models that leverages both conversation data across speakers and other types of data pertaining to the speaker and speaker roles to be modeled. Experiments show that our approach leads to significant improvements over baseline model quality, generating responses that capture more precisely speakers’ traits and speaking styles. The model offers the benefits of being algorithmically simple and easy to implement, and not relying on large quantities of data representing specific individual speakers. Yi Luan, Chris Brockett, William B. Dolan, Jianfeng Gao 0001, Michel Galley |
IJCNLP(1) | 5 |
| 2017 | Image-Grounded Conversations: Multimodal Context for Natural Question and Response GenerationabstractThe popularity of image sharing on social media and the engagement it creates between users reflect the important role that visual context plays in everyday conversations. We present a novel task, Image Grounded Conversations (IGC), in which natural-sounding conversations are generated about a shared image. To benchmark progress, we introduce a new multiple reference dataset of crowd-sourced, event-centric conversations on images. IGC falls on the continuum between chit-chat and goal-directed conversation models, where visual grounding constrains the topic of conversation to event-driven utterances. Experiments with models trained on social media data show that the combination of visual and textual context enhances the quality of generated conversational turns. In human evaluation, the gap between human performance and that of both neural and retrieval architectures suggests that multi-modal IGC presents an interesting challenge for dialog research. Nasrin Mostafazadeh, Chris Brockett, William B. Dolan, Michel Galley, Jianfeng Gao 0001, Georgios Spithourakis, Lucy Vanderwende |
IJCNLP(1) | 4 |
| 2016 | A Persona-Based Neural Conversation ModelabstractJiwei Li, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao, Bill Dolan. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Jiwei Li 0001, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao 0001, William B. Dolan |
ACL (1) | 2 |
| 2016 | Deep Reinforcement Learning for Dialogue GenerationabstractRecent neural models of dialogue generation offer great promise for generating responses for conversational agents, but tend to be shortsighted, predicting utterances one at a time while ignoring their influence on future outcomes.Modeling the future direction of a dialogue is crucial to generating coherent, interesting dialogues, a need which led traditional NLP models of dialogue to draw on reinforcement learning.In this paper, we show how to integrate these goals, applying deep reinforcement learning to model future reward in chatbot dialogue.The model simulates dialogues between two virtual agents, using policy gradient methods to reward sequences that display three useful conversational properties: informativity, coherence, and ease of answering (related to forward-looking function).We evaluate our model on diversity, length as well as with human judges, showing that the proposed algorithm generates more interactive responses and manages to foster a more sustained conversation in dialogue simulation.This work marks a first step towards learning a neural conversational model based on the long-term success of dialogues. Jiwei Li 0001, Will Monroe, Alan Ritter, Daniel Jurafsky, Michel Galley, Jianfeng Gao 0001 |
EMNLP | 5 |
| 2016 | Visual StorytellingabstractTing-Hao Kenneth Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, Margaret Mitchell. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Ting-Hao 'Kenneth' Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross B. Girshick, Xiaodong He 0001, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, Margaret Mitchell |
HLT-NAACL | 14 |
| 2016 | A Diversity-Promoting Objective Function for Neural Conversation ModelsabstractJiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, Bill Dolan. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Jiwei Li 0001, Michel Galley, Chris Brockett, Jianfeng Gao 0001, William B. Dolan |
HLT-NAACL | 2 |
| 2015 | Language to Code: Learning Semantic Parsers for If-This-Then-That RecipesabstractChris Quirk, Raymond Mooney, Michel Galley. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Chris Quirk, Raymond J. Mooney, Michel Galley |
ACL (1) | 3 |
| 2015 | A Survey of Current Datasets for Vision and Language ResearchabstractFrancis Ferraro, Nasrin Mostafazadeh, Ting-Hao Huang, Lucy Vanderwende, Jacob Devlin, Michel Galley, Margaret Mitchell. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Francis Ferraro, Nasrin Mostafazadeh, Ting-Hao 'Kenneth' Huang, Lucy Vanderwende, Jacob Devlin, Michel Galley, Margaret Mitchell |
EMNLP | 6 |
| 2015 | A Neural Network Approach to Context-Sensitive Generation of Conversational ResponsesabstractAlessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, Bill Dolan. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao 0001, William B. Dolan |
HLT-NAACL | 2 |
| 2014 | Large-scale Expected BLEU Training of Phrase-based Reordering ModelsabstractRecent work by Cherry (2013) has shown that directly optimizing phrase-based re-ordering models towards BLEU can lead to significant gains. Their approach is lim-ited to small training sets of a few thou-sand sentences and a similar number of sparse features. We show how the ex-pected BLEU objective allows us to train a simple linear discriminative reordering model with millions of sparse features on hundreds of thousands of sentences re-sulting in significant improvements. A comparison to likelihood training demon-strates that expected BLEU is vastly more effective. Our best results improve a hi-erarchical lexicalized reordering baseline by up to 2.0 BLEU in a single-reference setting on a French-English WMT 2012 setup. 1 Michael Auli, Michel Galley, Jianfeng Gao 0001 |
EMNLP | 2 |
| 2014 | AutoCaption: Automatic caption generation for personal photosabstractAutoCaption is a system that helps a smartphone user generate a caption for their photos. It operates by uploading the photo to a cloud service where a number of parallel modules are applied to recognize a variety of entities and relations. The outputs of the modules are combined to generate a large set of candidate captions, which are returned to the phone. The phone client includes a convenient user interface that allows users to select their favorite caption, reorder, add, or delete words to obtain the grammatical style they prefer. The user can also select from multiple candidates returned by the recognition modules. Krishnan Ramnath, Simon Baker, Lucy Vanderwende, Motaz Ahmad El-Saban, Sudipta N. Sinha, Anitha Kannan, Noran Hassan, Michel Galley, Yi Yang 0007, Deva Ramanan, Alessandro Bergamo, Lorenzo Torresani |
WACV | 8 |
| 2013 | Joint Language and Translation Modeling with Recurrent Neural NetworksabstractWe present a joint language and translation model based on a recurrent neural network which predicts target words based on an unbounded history of both source and target words.The weaker independence assumptions of this model result in a vastly larger search space compared to related feedforward-based language or translation models.We tackle this issue with a new lattice rescoring algorithm and demonstrate its effectiveness empirically.Our joint model builds on a well known recurrent neural network language model (Mikolov, 2012) augmented by a layer of additional inputs from the source language.We show competitive accuracy compared to the traditional channel model features.Our best results improve the output of a system trained on WMT 2012 French-English data by up to 1.5 BLEU, and by 1.1 BLEU on average across several test sets. Michael Auli, Michel Galley, Chris Quirk, Geoffrey Zweig |
EMNLP | 2 |
| 2013 | Regularized Minimum Error Rate TrainingabstractMinimum Error Rate Training (MERT) remains one of the preferred methods for tuning linear parameters in machine translation systems, yet it faces significant issues.First, MERT is an unregularized learner and is therefore prone to overfitting.Second, it is commonly used on a noisy, non-convex loss function that becomes more difficult to optimize as the number of parameters increases.To address these issues, we study the addition of a regularization term to the MERT objective function.Since standard regularizers such as ℓ 2 are inapplicable to MERT due to the scale invariance of its objective function, we turn to two regularizers-ℓ 0 and a modification of ℓ 2and present methods for efficiently integrating them during search.To improve search in large parameter spaces, we also present a new direction finding algorithm that uses the gradient of expected BLEU to orient MERT's exact line searches.Experiments with up to 3600 features show that these extensions of MERT yield results comparable to PRO, a learner often used with large feature sets. Michel Galley, Chris Quirk, Colin Cherry, Kristina Toutanova |
EMNLP | 1 |
| 2011 | Optimal Search for Minimum Error Rate Training
Michel Galley, Chris Quirk |
EMNLP | 1 |
| 2010 | Accurate Non-Hierarchical Phrase-Based Translation
Michel Galley, Christopher D. Manning |
HLT-NAACL | 1 |
| 2010 | Improved Models of Distortion Cost for Statistical Machine Translation
Spence Green, Michel Galley, Christopher D. Manning |
HLT-NAACL | 2 |
| 2009 | Quadratic-Time Dependency Parsing for Machine Translation
Michel Galley, Christopher D. Manning |
ACL/IJCNLP | 1 |
| 2009 | Robust Machine Translation Evaluation with Entailment Features
Sebastian Padó, Michel Galley, Daniel Jurafsky, Christopher D. Manning |
ACL/IJCNLP | 2 |
| 2009 | Measuring machine translation quality as semantic equivalence: A metric based on entailment features
Sebastian Padó, Daniel M. Cer, Michel Galley, Daniel Jurafsky, Christopher D. Manning |
Mach. Transl. | 3 |
| 2008 | A Simple and Effective Hierarchical Phrase Reordering Model
Michel Galley, Christopher D. Manning |
EMNLP | 1 |
| 2008 | A Phrase-Based Alignment Model for Natural Language Inference
Bill MacCartney, Michel Galley, Christopher D. Manning |
EMNLP | 2 |
| 2007 | Lexicalized Markov Grammars for Sentence Compression
Michel Galley, Kathy McKeown |
HLT-NAACL | 1 |
| 2006 | Automatic Summarization of Conversational Multi-Party Speech
Michel Galley |
AAAI | 1 |
| 2006 | Scalable Inference and Training of Context-Rich Syntactic Translation ModelsabstractStatistical MT has made great progress in the last few years, but current translation models are weak on re-ordering and target language fluency. Syntactic approaches seek to remedy these problems. In this paper, we take the framework for acquiring multi-level syntactic translation rules of (Galley et al., 2004) from aligned tree-string pairs, and present two main extensions of their approach: first, instead of merely computing a single derivation that minimally explains a sentence pair, we construct a large number of derivations that include contextually richer rules, and account for multiple interpretations of unaligned words. Second, we propose probability estimates and a training procedure for weighting these rules. We contrast different approaches on real examples, show that our estimates based on multiple derivations favor phrasal re-orderings that are linguistically better motivated, and establish that our larger rules provide a 3.63 BLEU point increase over minimal rules. Michel Galley, Jonathan Graehl, Kevin Knight, Daniel Marcu, Steve DeNeefe, Wei Wang 0006, Ignacio Thayer |
ACL | 1 |
| 2006 | A Skip-Chain Conditional Random Field for Ranking Meeting Utterances by Importance
Michel Galley |
EMNLP | 1 |
| 2005 | From text to speech summarizationabstractIn this paper, we present approaches used in text summarization, showing how they can be adapted for speech summarization and where they fall short. Informal style and apparent lack of structure in speech mean that the typical approaches used for text summarization must be extended for use with speech. We illustrate how features derived from speech can help determine summary content within two ongoing summarization projects at Columbia University. Kathy McKeown, Julia Hirschberg, Michel Galley, Sameer Maskey |
ICASSP (5) | 3 |
| 2004 | Identifying Agreement and Disagreement in Conversational Speech: Use of Bayesian Networks to Model Pragmatic DependenciesabstractWe describe a statistical approach for modeling agreements and disagreements in conversational interaction. Our approach first identifies adjacency pairs using maximum entropy ranking based on a set of lexical, durational, and structural features that look both forward and backward in the discourse. We then classify utterances as agreement or disagreement using these adjacency pairs and features that represent various pragmatic influences of previous agreement or disagreement on the current utterance. Our approach achieves 86.9% accuracy, a 4.9% increase over previous work. Michel Galley, Kathy McKeown, Julia Hirschberg, Elizabeth Shriberg |
ACL | 1 |
| 2004 | What's in a translation rule?
Michel Galley, Mark Hopkins, Kevin Knight, Daniel Marcu |
HLT-NAACL | 1 |
| 2003 | Discourse Segmentation of Multi-Party ConversationabstractWe present a domain-independent topic segmentation algorithm for multi-party speech. Our feature-based algorithm combines knowledge about content using a text-based algorithm as a feature and about form using linguistic and acoustic cues about topic shifts extracted from speech. This segmentation algorithm uses automatically induced decision rules to combine the different features. The embedded text-based algorithm builds on lexical cohesion and has performance comparable to state-of-the-art algorithms based on lexical information. A significant error reduction is obtained by combining the two knowledge sources. Michel Galley, Kathy McKeown, Eric Fosler-Lussier, Hongyan Jing |
ACL | 1 |
| 2003 | Improving Word Sense Disambiguation in Lexical Chaining
Michel Galley, Kathy McKeown |
IJCAI | 1 |
| 2001 | Hybrid natural language generation for spoken dialogue systemsabstractThe natural language generation component of most dialogue systems is based on templates. Template-based generators are hard to maintain and reuse, and the sentences they produce lack the variability and robustness needed by conversational systems. In this paper, a flexible and domain-independent natural language generator for spoken dialogue systems is proposed which combines fixed surface expressions with freely generated text. The generation algorithm follows a hybrid approach, combining finite state machine (FSM) grammars and corpus-based language models. In this approach, the FSM grammar (a reversible parser grammar) is constrained by a word and concept Ò-gram that takes terminals and non-terminal co-occurrences into account. The Ò-gram grammar helps prevent inappropriate derivations, therefore improving the quality of the generated texts. The proposed algorithm achieves faster than real-time performance because of the limited number of derivations. Michel Galley, Eric Fosler-Lussier, Alexandros Potamianos |
INTERSPEECH | 1 |