EDBT 2026 Demo / reviewers in the wild / expert
Zhou Yu 0005
dblp:83/3205-5
· DBLP profile ↗
74ranked-venue papers
6as first author
46since 2021 · last 2026
0000-0002-1524-5890ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 69 · 6 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic Volume: Quantifying and Detecting Both External and Internal Uncertainty in LLMsabstractLarge language models (LLMs) have demonstrated remarkable performance across diverse tasks by encoding vast amounts of factual knowledge. However, they are still prone to hallucinations, generating incorrect or misleading information, often accompanied by high uncertainty. Existing methods for hallucination detection primarily focus on quantifying internal uncertainty, which arises from missing or conflicting knowledge within the model. However, hallucinations can also stem from external uncertainty, where ambiguous user queries lead to multiple possible interpretations. In this work, we introduce **Semantic Volume**, a novel mathematical measure for quantifying both external and internal uncertainty in LLMs. Our approach perturbs queries and responses, embeds them in a semantic space, and computes the Gram matrix determinant of the embedding vectors, capturing their dispersion as a measure of uncertainty. Our framework provides a generalizable and unsupervised uncertainty detection method without requiring internal access to LLMs. We conduct extensive experiments on both external and internal uncertainty detections, demonstrating that our Semantic Volume method consistently outperforms existing baselines in both tasks. Additionally, we provide theoretical insights linking our measure to differential entropy, unifying and extending previous sampling-based uncertainty measures such as the semantic entropy. Semantic Volume is shown to be a robust and interpretable approach to improving the reliability of LLMs by systematically detecting uncertainty in both user queries and model responses. Zhou Yu 0005, Ziji Zhang 0003, Yingying Zhuang, Swair Shah, Narayanan Sadagopan, Anurag Beniwal |
AAAI | 2 |
| 2026 | DXA-Net: Dual-Task Cross-Lingual Alignment Network for Zero-Shot Cross-Lingual Spoken Language UnderstandingabstractThe state-of-the-art zero-shot cross-lingual spoken language understanding (SLU) model utilizes cross-lingual unsupervised contrastive learning to achieve multilingual semantics alignment. While existing methods have achieved promising results, they still have two issues limiting cross-lingual knowledge transfer: (1) dual-task correlative knowledge is not explicitly modeled and transferred to target languages; (2) the semantics differences among samples are ignored, and the contrastive semantics knowledge is not transferred to target languages. In this paper, we propose a dual-task cross-lingual alignment network (DXA-Net), which makes the first attempt to tackle zero-shot cross-lingual SLU based on the prompt-tuning paradigm. To solve the first issue, we propose the co-guiding prompt, which allows the model to conditionally generate one task's label based on another one's. To solve the second issue, we propose the intent/slot contrastive prompt to teach the model to discriminate whether a pair of samples have the same or similar labels. Additionally, we propose multilingual semantics contrastive prompt to enhance multilingual semantics alignment. Experiments on the benchmark show that our model achieves new state-of-the-art performance on nine languages. Libo Qin 0001, Zhihong Zhu 0001, Zhou Yu 0005, Ivor W. Tsang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory LearningabstractAutonomous agents have demonstrated significant potential in automating complex multistep decision-making tasks. However, even state-of-the-art vision-language models (VLMs), such as GPT-4o, still fall short of human-level performance, particularly in intricate web environments and long-horizon planning tasks. To address these limitations, we introduce Reflective Monte Carlo Tree Search (R-MCTS), a novel test-time algorithm designed to enhance the ability of AI agents, e.g., powered by GPT-4o, to explore decision space on the fly.
R-MCTS extends traditional MCTS by 1) incorporating contrastive reflection, allowing agents to learn from past interactions and dynamically improve their search efficiency; and 2) using multi-agent debate to provide reliable state evaluation. Moreover, we improve the agent's performance by fine-tuning GPT-4o through self-learning, using R-MCTS generated tree traversals without any human-provided labels. On the challenging VisualWebArena benchmark, our GPT-4o-based R-MCTS agent achieves a 6% to 30% relative improvement across various tasks compared to the previous state-of-the-art. Additionally, we show that the knowledge gained from test-time search can be effectively transferred back to GPT-4o via fine-tuning. The fine-tuned GPT-4o matches 97\% of R-MCTS's performance while reducing compute usage by a factor of four at test time. Furthermore, qualitative results reveal that the fine-tuned GPT-4o model demonstrates the ability to explore the environment, evaluate a state, and backtrack to viable ones when it detects that the current state cannot lead to success. Moreover, our work demonstrates the compute scaling properties in both training - data collection with R-MCTS - and testing time. These results suggest a promising research direction to enhance VLMs' reasoning and planning capabilities for agentic applications via test-time search and self-learning. Xiao Yu 0011, Baolin Peng, Vineeth Vajipey, Hao Cheng 0002, Michel Galley, Jianfeng Gao 0001, Zhou Yu 0005 |
ICLR | 7 |
| 2025 | When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMsabstractReasoning-enhanced large language models (RLLMs), whether explicitly trained for reasoning or prompted via chain-of-thought (CoT), have achieved state-of-the-art performance on many complex reasoning tasks. However, we uncover a surprising and previously overlooked phenomenon: explicit CoT reasoning can significantly degrade instruction-following accuracy. Evaluating 20+ models on two benchmarks: IFEval (with simple, rule-verifiable constraints) and ComplexBench (with complex, compositional constraints), we consistently observe performance drops when CoT prompting is applied. Through large-scale case studies and an attention-based analysis, we identify common patterns where reasoning either helps (e.g., with formatting or lexical precision) or hurts (e.g., by neglecting simple constraints or introducing unnecessary content). We propose a metric, constraint attention, to quantify model focus during generation and show that CoT reasoning often diverts attention away from instruction-relevant tokens. To mitigate these effects, we introduce and evaluate four strategies: in-context learning, self-reflection, self-selective reasoning, and classifier-selective reasoning. Our results demonstrate that selective reasoning strategies, particularly classifier-selective reasoning, can substantially recover lost performance. To our knowledge, this is the first work to systematically expose reasoning-induced failures in instruction-following and offer practical mitigation strategies. Zhou Yu 0005, Xupeng Chen, Ziji Zhang 0003, Yingying Zhuang, Narayanan Sadagopan, Anurag Beniwal |
NeurIPS | 2 |
| 2024 | The Mirrored Influence Hypothesis: Efficient Data Influence Estimation by Harnessing Forward PassesabstractLarge-scale black-box models have become ubiquitous across numerous applications. Understanding the influence of individual training data sources on predictions made by these models is crucial for improving their trustworthiness. Current influence estimation techniques involve computing gradients for every training point or repeated training on different subsets. These approaches face obvious computational challenges when scaled up to large datasets and models. In this paper, we introduce and explore the Mirrored Influence Hypothesis, highlighting a reciprocal nature of influence between training and test data. Specifically, it sug-gests that evaluating the influence of training data on test predictions can be reformulated as an equivalent, yet inverse problem: assessing how the predictions for training samples would be altered if the model were trained on specific test samples. Through both empirical and theoretical validations, we demonstrate the wide applicability of our hypothesis. Inspired by this, we introduce a new method for estimating the influence of training data, which requires calculating gradients for specific test samples, paired with a forward pass for each training point. This approach can capitalize on the common asymmetry in scenarios where the number of test samples under concurrent examination is much smaller than the scale of the training dataset, thus gaining a significant improvement in efficiency compared to existing approaches. We demonstrate the applicability of our method across a range of scenarios, including data attribution in diffusion models, data leakage detection, analy-sis of memorization, mislabeled data detection, and tracing behavior in language models. Myeongseob Ko, Feiyang Kang, Weiyan Shi 0001, Ming Jin 0002, Zhou Yu 0005, Ruoxi Jia 0001 |
CVPR | 5 |
| 2024 | LIONs: An Empirically Optimized Approach to Align Language ModelsabstractAlignment is a crucial step to enhance the instruction-following and conversational abilities of language models.Despite many recent work proposing new algorithms, datasets, and training pipelines, there is a lack of comprehensive studies measuring the impact of various design choices throughout the whole training process.We first conduct a rigorous analysis over a three-stage training pipeline consisting of supervised fine-tuning, offline preference learning, and online preference learning.We have found that using techniques like sequence packing, loss masking in SFT, increasing the preference dataset size in DPO, and online DPO training can significantly improve the performance of language models.We then train from Gemma-2b-base and LLama-3-8b-base, and find that our best models exceed the performance of the official instruct models tuned with closed-source data and algorithms.Our code and models can be found at https://github. Xiao Yu 0011, Qingyang Wu, Yu Li 0013, Zhou Yu 0005 |
EMNLP | 4 |
| 2024 | DECOR: Improving Coherence in L2 English Writing with a Novel Benchmark for Incoherence Detection, Reasoning, and RewritingabstractCoherence in writing, an aspect that secondlanguage (L2) English learners often struggle with, is crucial in assessing L2 English writing.Existing automated writing evaluation systems primarily use basic surface linguistic features to detect coherence in writing.However, little effort has been made to correct the detected incoherence, which could significantly benefit L2 language learners seeking to improve their writing.To bridge this gap, we introduce DECOR, a novel benchmark that includes expert annotations for detecting incoherence in L2 English writing, identifying the underlying reasons, and rewriting the incoherent sentences.To our knowledge, DECOR is the first coherence assessment dataset specifically designed for improving L2 English writing, featuring pairs of original incoherent sentences alongside their expert-rewritten counterparts.Additionally, we fine-tuned models to automatically detect and rewrite incoherence in student essays.We find that incorporating specific reasons for incoherence during fine-tuning consistently improves the quality of the rewrites, achieving a result that is favored in both automatic and human evaluations.1 Xuanming Zhang, Anthony Diaz, Zixun Chen, Qingyang Wu, Kun Qian 0016, Erik Voss, Zhou Yu 0005 |
EMNLP | 7 |
| 2024 | MIRACLE: An Online, Explainable Multimodal Interactive Concept Learning SystemabstractWe present MIRACLE, a system for online, interpretable visual concept and video action recognition. Through a chat interface, users query the recognition system with an uploaded image or video. For images, MIRACLE returns concept predictions from its structured knowledge base, justifying its predictions with heatmaps and natural language-based attribute detections. For videos, MIRACLE predicts an action and justifies its prediction with time varying entity-entity relations. With its ability to learn new concepts in an online, few-shot manner and its support of dynamic changes to its knowledge base, MIRACLE represents a step forward in interpretable multimodal learning systems. Ansel Blume, Khanh Duy Nguyen, Zhenhailong Wang, Yangyi Chen, Michal Shlapentokh-Rothman, Xiaomeng Jin, Zhen Zhu 0006, Jiateng Liu, Kuan-Hao Huang, Mankeerat Sidhu, Xuanming Zhang, Vivian Liu, Raunak Sinha, Te-Lin Wu, Abhaysinh Zala, Elias Stengel-Eskin, Da Yin, Utkarsh Mall, Zhou Yu 0005, Kai-Wei Chang 0001, Camille Cobb, Karrie Karahalios, Lydia B. Chilton, Mohit Bansal, Nanyun Peng 0001, Carl Vondrick, Derek Hoiem, Heng Ji 0001 |
ACM Multimedia | 21 |
| 2024 | Teaching Language Models to Self-Improve through Interactive DemonstrationsabstractXiao Yu, Baolin Peng, Michel Galley, Jianfeng Gao, Zhou Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xiao Yu 0011, Baolin Peng, Michel Galley, Jianfeng Gao 0001, Zhou Yu 0005 |
NAACL-HLT | 5 |
| 2024 | ConFit: Improving Resume-Job Matching using Data Augmentation and Contrastive LearningabstractA reliable resume-job matching system helps a company find suitable candidates from a pool of resumes, and helps a job seeker find relevant jobs from a list of job posts. However, since job seekers apply only to a few jobs, interaction records in resume-job datasets are sparse. Different from many prior work that use complex modeling techniques, we tackle this sparsity problem using data augmentations and a simple contrastive learning approach. ConFit first formulates resume-job datasets as a sparse bipartite graph, and creates an augmented dataset by paraphrasing specific sections in a resume or a job post. Then, ConFit finetunes pre-trained encoders with contrastive learning to further increase training samples from B pairs per batch to <?TeX $\mathcal {O}(B^2)$?> Math 1 per batch. We evaluate ConFit on two real-world datasets and find it outperforms prior methods (including BM25 and OpenAI text-ada-002) by up to 19% and 31% absolute in nDCG@10 for ranking jobs and ranking resumes, respectively. We believe ConFit’s simple yet highly performant approach lays a strong foundation for future research in modeling person-job fit.1 Xiao Yu 0011, Jinzhong Zhang 0002, Zhou Yu 0005 |
RecSys | 3 |
| 2024 | Curriculum-Driven Edubot: A Framework for Developing Language Learning Chatbots through Synthesizing Conversational DataabstractChatbots have become popular in educational settings, revolutionizing how students interact with material and how teachers teach.We present Curriculum-Driven EduBot, a framework for developing a chatbot that combines the interactive features of chatbots with the systematic material of English textbooks to assist students in enhancing their conversational skills.We begin by extracting pertinent topics from textbooks and using large language models to generate dialogues related to these topics.We then fine-tune an open-source model using our generated conversational data to create our curriculum-driven chatbot.User studies demonstrate that EduBot outperforms ChatGPT in leading curriculum-based dialogues and adapting its dialogue to match the user's English proficiency level.By combining traditional textbook methodologies with conversational AI, our approach offers learners an interactive tool that aligns with their curriculum and provides user-tailored conversation practice.This facilitates meaningful student-bot dialogues and enriches the overall learning experience within the curriculum's pedagogical framework. Yu Li 0013, Shang Qu, Jili Shen, Shangchao Min, Zhou Yu 0005 |
SIGDIAL | 5 |
| 2024 | Dialoging Resonance in Human-Chatbot Conversation: How Users Perceive and Reciprocate Recommendation Chatbot's Self-Disclosure StrategyabstractUsing chatbots to make recommendations is increasingly popular. The design of recommendation chatbots has mainly been taking an information-centric approach by focusing on the recommended content per se. Limited attention is on how social connection and relational strategies, such as self-disclosure from a chatbot, may influence users' perception and acceptance of the recommendation. In this work, we designed, implemented, and evaluated a social chatbot capable of performing three different levels of self-disclosure: factual information (low), cognitive opinions (medium), and emotions (high). In the evaluation, we recruited 372 participants to converse with the chatbot on two topics: movies and COVID-19 experiences. In each topic, the chatbot conducted small talks and made relevant recommendations to the topic. Participants were randomly assigned to four experimental conditions where the chatbot used factual, cognitive, emotional, and adaptive strategies to perform self-disclosures. By training a text classifier to identify users' level of self-disclosure in real-time, the adaptive chatbot can dynamically match its self-disclosure language to the level of disclosure exhibited by the users. Our results show that users reciprocate with higher-level self-disclosure when a recommendation chatbot displays emotions throughout the conversation. The utilization of emotional disclosure by the chatbot resulted in enhanced enjoyment during interactions and a more favorable perception of the bot. This, in turn, led to greater effectiveness in making recommendations, including a higher likelihood of accepting the recommendation. We discuss the understandings obtained and implications to future design. Kaihui Liang, Weiyan Shi 0001, Hao-Chuan Wang, Zhou Yu 0005 |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2023 | DIONYSUS: A Pre-trained Model for Low-Resource Dialogue SummarizationabstractDialogue summarization has recently garnered significant attention due to its wide range of applications.However, existing methods for summarizing dialogues have limitations because they do not take into account the inherent structure of dialogue and rely heavily on labeled data, which can lead to poor performance in new domains.In this work, we propose DIONYSUS (dynamic input optimization in pre-training for dialogue summarization), a pre-trained encoder-decoder model for summarizing dialogues in any new domain.To pretrain DIONYSUS, we create two pseudo summaries for each dialogue example: one from a fine-tuned summarization model and the other from important dialogue turns.We then choose one of these pseudo summaries based on information distribution differences in different types of dialogues.This selected pseudo summary serves as the objective for pre-training DIONYSUS using a self-supervised approach on a large dialogue corpus.Our experiments show that DIONYSUS outperforms existing methods on six datasets, as demonstrated by its ROUGE scores in zero-shot and few-shot settings. Yu Li 0013, Baolin Peng, Michel Galley, Zhou Yu 0005, Jianfeng Gao 0001 |
ACL (1) | 5 |
| 2023 | State Value Generation with Prompt Learning and Self-Training for Low-Resource Dialogue State Tracking
Yan Yang 0008, Chengcai Chen, Zhou Yu 0005 |
ACML | 4 |
| 2023 | Social Influence Dialogue Systems: A Survey of Datasets and Models For Social Influence TasksabstractDialogue systems capable of social influence such as persuasion, negotiation, and therapy, are essential for extending the use of technology to numerous realistic scenarios.However, existing research primarily focuses on either task-oriented or open-domain scenarios, a categorization that has been inadequate for capturing influence skills systematically.There exists no formal definition or category for dialogue systems with these skills and data-driven efforts in this direction are highly limited.In this work, we formally define and introduce the category of social influence dialogue systems that influence users' cognitive and emotional responses, leading to changes in thoughts, opinions, and behaviors through natural conversations.We present a survey of various tasks, datasets, and methods, compiling the progress across seven diverse domains.We discuss the commonalities and differences between the examined systems, identify limitations, and recommend future directions.This study serves as a comprehensive reference for social influence dialogue systems to inspire more dedicated research and discussion in this emerging area. Kushal Chawla, Weiyan Shi 0001, Gale M. Lucas, Zhou Yu 0005, Jonathan Gratch |
EACL | 5 |
| 2023 | FastKASSIM: A Fast Tree Kernel-Based Syntactic Similarity MetricabstractSyntax is a fundamental component of language, yet few metrics have been employed to capture syntactic similarity or coherence at the utterance-and document-level.The existing standard document-level syntactic similarity metric is computationally expensive and performs inconsistently when faced with syntactically dissimilar documents.To address these challenges, we present FastKASSIM, a metric for utterance-and document-level syntactic similarity which pairs and averages the most similar constituency parse trees between a pair of documents based on tree kernels.FastKAS-SIM is more robust to syntactic dissimilarities and runs up to to 5.32 times faster than its predecessor over documents in the r/ChangeMyView corpus.FastKASSIM's improvements allow us to examine hypotheses in two settings with large documents.We find that syntactically similar arguments on r/ChangeMyView tend to be more persuasive, and that syntax is predictive of authorship attribution in the Australian High Court Judgment corpus. * denotes equal contribution.Utterance 1: When we hate, we always move away from the grace of God.When we become resentful and unforgiving, the world around us seems spiteful and meaningless.Utterance 2: How can you be skiing if you are already swimming?FastKASSIM Score: 0.219 CASSIM Score: 0.838 LSM Score: 0.623 Utterance 1: I like swimming because it is cool.Utterance 2: I love running because it is fun. Maximillian Chen 0001, Caitlyn Chen, Xiao Yu 0011, Zhou Yu 0005 |
EACL | 4 |
| 2023 | End-to-end Task-oriented Dialogue: A Survey of Tasks, Methods, and Future DirectionsabstractEnd-to-end task-oriented dialogue (EToD) can directly generate responses in an end-to-end fashion without modular training, which attracts escalating popularity.The advancement of deep neural networks, especially the successful use of large pre-trained models, has further led to significant progress in EToD research in recent years.In this paper, we present a thorough review and provide a unified perspective to summarize existing approaches as well as recent trends to advance the development of EToD research.The contributions of this paper can be summarized: (1) First survey: to our knowledge, we take the first step to present a thorough survey of this research field; (2) New taxonomy: we first introduce a unified perspective for EToD, including (i) Modularly EToD and (ii) Fully EToD; (3) New Frontiers: we discuss some potential frontier areas as well as the corresponding challenges, hoping to spur breakthrough research in EToD field; (4) Abundant resources: we build a public website 1 , where EToD researchers could directly access the recent progress.We hope this work can serve as a thorough reference for the EToD research community.EToD Modularly EToD ( §3.1) Libo Qin 0001, Wenbo Pan 0001, Qiguang Chen, Lizi Liao, Zhou Yu 0005, Yue Zhang 0004, Wanxiang Che, Min Li 0007 |
EMNLP | 5 |
| 2023 | Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy PlanningabstractPlanning for goal-oriented dialogue often requires simulating future dialogue interactions and estimating task progress.Many approaches thus consider training neural networks to perform look-ahead search algorithms such as A* search and Monte Carlo Tree Search (MCTS).However, this training often requires abundant annotated data, which creates challenges when faced with noisy annotations or low-resource settings.We introduce GDP-ZERO, an approach using Open-Loop MCTS to perform goal-oriented dialogue policy planning without any model training.GDP-ZERO prompts a large language model to act as a policy prior, value function, user simulator, and system model during the tree search.We evaluate GDP-ZERO on the goal-oriented task Persua-sionForGood, and find that its responses are preferred over ChatGPT up to 59.32% of the time, and are rated more persuasive than Chat-GPT during interactive evaluations 1 . Xiao Yu 0011, Maximillian Chen 0001, Zhou Yu 0005 |
EMNLP | 3 |
| 2023 | KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords LearningabstractIn task-oriented dialogs (TOD), reinforcement learning (RL) algorithms train a model to directly optimize response for task-related metrics.However, RL needs to perform exploration, which can be time-consuming due to the slow auto-regressive sequence generation process.We investigate an approach to create a more efficient RL-based algorithm to improve TOD performance in an offline setting.First, we use a faster generation procedure that samples from independent next-word distributions after training the language model (LM) with supervised learning.We then introduce a finegrained reward function to help the model focus on learning key information in a dialog, by measuring the importance and semantic closeness of each generated token.Experiments on the MultiWoZ dataset show our new training algorithm, Keywords Reinforcement Learning with Next-word Sampling (KRLS), achieves state-of-the-art performance on the end-to-end response generation task, with a 15% training time reduction compared to a standard RL algorithm using auto-regressive generation 1 . Xiao Yu 0011, Qingyang Wu, Kun Qian 0016, Zhou Yu 0005 |
EMNLP | 4 |
| 2023 | Pre-Finetuning for Few-Shot Emotional Speech RecognitionabstractSpeech models have long been known to overfit individual speakers for many classification tasks.This leads to poor generalization in settings where the speakers are out-of-domain or out-of-distribution, as is common in production environments.We view speaker adaptation as a few-shot learning problem and propose investigating transfer learning approaches inspired by recent success with pre-trained models in natural language tasks.We propose pre-finetuning speech models on difficult tasks to distill knowledge into few-shot downstream classification objectives.We pre-finetune Wav2Vec2.0 on every permutation of four multiclass emotional speech recognition corpora and evaluate our pre-finetuned models through 33,600 few-shot fine-tuning trials on the Emotional Speech Dataset. Maximillian Chen 0001, Zhou Yu 0005 |
INTERSPEECH | 2 |
| 2023 | Learning Joint Policies for Human-Robot Dialog and Co-NavigationabstractService robots need language capabilities for communicating with people, and navigation skills for beyond-proximity interaction in the real world. When the robot explores the real world with people side by side, there is the compound problem of human-robot dialog and co-navigation. The human-robot team uses dialog to decide where to go, and their shared spatial awareness affects the dialog state. In this paper, we develop a framework that learns a joint policy for human-robot dialog and co-navigation toward efficiently and accurately completing tour guide and information delivery tasks. We show that our approach outperforms baselines from the literature in task completion rate and execution time, and demonstrate our approach in the real world. Yohei Hayamizu, Zhou Yu 0005, Shiqi Zhang 0001 |
IROS | 2 |
| 2023 | Towards Next-Generation Intelligent Assistants Leveraging LLM TechniquesabstractVirtual Intelligent Assistants take user requests in the voice form, perform actions such as setting an alarm, turning on a light, and answering a question, and provide answers or confirmations in the voice form or through other channels such as a screen. Assistants have become prevalent in the past decade, and users have been taking services from assistants like Amazon Alexa, Apple Siri, Google Assistant, and Microsoft Cortana. Xin Dong 0001, Seungwhan Moon, Yifan Ethan Xu, Kshitiz Malik, Zhou Yu 0005 |
KDD | 5 |
| 2022 | Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative ComprehensionabstractYing Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Bingsheng Yao, Tongshuang Wu, Zheng Zhang, Toby Li, Nora Bradford, Branda Sun, Tran Hoang, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu, Mark Warschauer. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Dakuo Wang, Mo Yu, Daniel Ritchie 0002, Bingsheng Yao, Sherry Tongshuang Wu, Zheng Zhang 0043, Toby Jia-Jun Li, Nora Bradford, Branda Sun, Tran Bao Hoang, Yisi Sang, Yufang Hou 0001, Xiaojuan Ma, Diyi Yang, Nanyun Peng 0001, Zhou Yu 0005, Mark Warschauer |
ACL (1) | 17 |
| 2022 | CGIM: A Cycle Guided Interactive Learning Model for Consistency Identification in Task-oriented DialogueabstractConsistency identification in task-oriented dialog (CI-ToD) usually consists of three subtasks, aiming to identify inconsistency between current system response and current user response, dialog history and the corresponding knowledge base. This work aims to solve CI-ToD task by introducing an explicit interaction paradigm, Cycle Guided Interactive learning Model (CGIM), which achieves to make information exchange explicitly from all the three tasks. Specifically, CGIM relies on two core insights, referred to as guided multi-head attention module and cycle interactive mechanism, that collaborate from each other. On the one hand, each two tasks are linked with the guided multi-head attention module, aiming to explicitly model the interaction across two related tasks. On the other hand, we further introduce cycle interactive mechanism that focuses on facilitating model to exchange information among the three correlated sub-tasks via a cycle interaction manner. Experimental results on CI-ToD benchmark show that our model achieves the state-of-the-art performance, pushing the overall score to 56.3% (5.0% point absolute improvement). In addition, we find that CGIM is robust to the initial task flow order. Libo Qin 0001, Qiguang Chen, Tianbao Xie, Qian Liu 0033, Shijue Huang, Wanxiang Che, Zhou Yu 0005 |
COLING | 7 |
| 2022 | Unsupervised Vision-and-Language Pretraining via Retrieval-based Multi-Granular AlignmentabstractVision-and-Language (V+L) pre-training models have achieved tremendous success in recent years on various multi-modal benchmarks. However, the majority of existing models require pre-training on a large set of parallel imagetext data, which is costly to collect, compared to image-only or text-only data. In this paper, we explore unsupervised Vision-and-Language pre-training (UVLP) to learn the cross-modal representation from non-parallel image and text datasets. We found two key factors that lead to good unsupervised V + L pre-training without parallel data: (i) joint image-and-text input (ii) overall imagetext alignment (even for non-parallel data). Accordingly, we propose a novel unsupervised V + L pre-training curriculum for non-parallel texts and images. We first construct a weakly aligned imagetext corpus via a retrieval-based approach, then apply a set of multi-granular alignment pre-training tasks, including region-to-tag, region-to-phrase, and image-to-sentence alignment, to bridge the gap between the two modalities. A comprehensive ablation study shows each granularity is helpful to learn a stronger pre-trained model. We adapt our pre-trained model to a set of V+L downstream tasks, including VQA, NLVR2, Visual Entailment, and Ref-COCO+. Our model achieves the state-of-art performance in all these tasks under the unsupervised setting. Mingyang Zhou 0004, Licheng Yu, Amanpreet Singh, Mengjiao Wang 0002, Zhou Yu 0005 |
CVPR | 5 |
| 2022 | Robots-Dont-Cry: Understanding Falsely Anthropomorphic Utterances in Dialog SystemsabstractDialog systems are often designed or trained to output human-like responses.However, some responses may be impossible for a machine to truthfully say (e.g."that movie made me cry").Highly anthropomorphic responses might make users uncomfortable or implicitly deceive them into thinking they are interacting with a human.We collect human ratings on the feasibility of approximately 900 two-turn dialogs sampled from 9 diverse data sources.Ratings are for two hypothetical machine embodiments: a futuristic humanoid robot and a digital assistant.We find that for some data-sources commonly used to train dialog systems, 20-30% of utterances are not viewed as possible for a machine.Rating is marginally affected by machine embodiment.We explore qualitative and quantitative reasons for these ratings.Finally, we build classifiers and explore how modeling configuration might affect output permissibly, and discuss implications for building less falsely anthropomorphic dialog systems. David Gros 0001, Yu Li 0013, Zhou Yu 0005 |
EMNLP | 3 |
| 2022 | Just Fine-tune Twice: Selective Differential Privacy for Large Language ModelsabstractProtecting large language models from privacy leakage is becoming increasingly crucial with their wide adoption in real-world products.Yet applying differential privacy (DP), a canonical notion with provable privacy guarantees for machine learning models, to those models remains challenging due to the trade-off between model utility and privacy loss.Utilizing the fact that sensitive information in language data tends to be sparse, Shi et al. (2021) formalized a DP notion extension called Selective Differential Privacy (SDP) to protect only the sensitive tokens defined by a policy function.However, their algorithm only works for RNN-based models.In this paper, we develop a novel framework, Just Fine-tune Twice (JFT), that achieves SDP for state-of-the-art large transformer-based models.Our method is easy to implement: it first finetunes the model with redacted in-domain data, and then fine-tunes it again with the original in-domain data using a private training mechanism.Furthermore, we study the scenario of imperfect implementation of policy functions that misses sensitive tokens and develop systematic methods to handle it.Experiments show that our method achieves strong utility compared to previous baselines.We also analyze the SDP privacy guarantee empirically with the canary insertion attack 1 . Weiyan Shi 0001, Ryan Shea, Si Chen 0008, Chiyuan Zhang, Ruoxi Jia 0001, Zhou Yu 0005 |
EMNLP | 6 |
| 2022 | Using Chatbots to Teach LanguagesabstractThis paper reports on progress towards building an online language learning tool to provide learners with conversational experience by using dialog systems as conversation practice partners. Our system can adapt to users' language proficiency on the fly. We also provide automatic grammar error feedback to help users learn from their mistakes. According to our first adopters, our system is entertaining and useful. Furthermore, we will provide the learning technology community a large-scale conversation dataset on language learning and grammar correction. Our next step is to make our system more adaptive to user profile information by using reinforcement learning algorithms. Yu Li 0013, Dian Yu 0002, Sam Davidson, Ryan Hou, Xun Yuan 0002, Yinghua Tan, Derek Pham, Zhou Yu 0005 |
L@S | 9 |
| 2022 | ErAConD: Error Annotated Conversational Dialog Dataset for Grammatical Error CorrectionabstractCurrently available grammatical error correction (GEC) datasets are compiled using essays or other long-form text written by language learners, limiting the applicability of these datasets to other domains such as informal writing and conversational dialog.In this paper, we present a novel GEC dataset consisting of parallel original and corrected utterances drawn from open-domain chatbot conversations; this dataset is, to our knowledge, the first GEC dataset targeted to a human-machine conversational setting.We also present a detailed annotation scheme which ranks errors by perceived impact on comprehension, making our dataset more representative of realworld language learning applications.To demonstrate the utility of the dataset, we use our annotated data to fine-tune a state-of-theart GEC model.Experimental results show the effectiveness of our data in improving GEC model performance in a conversational scenario. Xun Yuan 0002, Derek Pham, Sam Davidson, Zhou Yu 0005 |
NAACL-HLT | 4 |
| 2022 | Knowledge-Grounded Dialogue Generation with a Unified Knowledge RepresentationabstractYu Li, Baolin Peng, Yelong Shen, Yi Mao, Lars Liden, Zhou Yu, Jianfeng Gao. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yu Li 0013, Baolin Peng, Yelong Shen, Lars Liden, Zhou Yu 0005, Jianfeng Gao 0001 |
NAACL-HLT | 6 |
| 2022 | Database Search Results Disambiguation for Task-Oriented Dialog SystemsabstractKun Qian, Satwik Kottur, Ahmad Beirami, Shahin Shayandeh, Paul Crook, Alborz Geramifard, Zhou Yu, Chinnadhurai Sankar. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Kun Qian 0016, Satwik Kottur, Ahmad Beirami, Shahin Shayandeh, Paul A. Crook, Alborz Geramifard, Zhou Yu 0005, Chinnadhurai Sankar |
NAACL-HLT | 7 |
| 2022 | Selective Differential Privacy for Language ModelingabstractWith the increasing applications of language models, it has become crucial to protect these models from leaking private information.Previous work has attempted to tackle this challenge by training RNN-based language models with differential privacy guarantees.However, applying classical differential privacy to language models leads to poor model performance as the underlying privacy notion is over-pessimistic and provides undifferentiated protection for all tokens in the data.Given that the private information in natural language is sparse (for example, the bulk of an email might not carry personally identifiable information), we propose a new privacy notion, selective differential privacy, to provide rigorous privacy guarantees on the sensitive portion of the data to improve model utility.To realize such a new notion, we develop a corresponding privacy mechanism, Selective-DPSGD, for RNN-based language models.Besides language modeling, we also apply the method to a more concrete application -dialog systems.Experiments on both language modeling and dialog system building show that the proposed privacy-preserving mechanism achieves better utilities while remaining safe under various privacy attacks compared to the baselines.The data and code are released to facilitate future research 1 . Weiyan Shi 0001, Aiqi Cui, Evan Li, Ruoxi Jia 0001, Zhou Yu 0005 |
NAACL-HLT | 5 |
| 2022 | Towards Socially Intelligent Agents with Mental State Transition and Human ValueabstractBuilding a socially intelligent agent involves many challenges.One of which is to track the agent's mental state transition and teach the agent to make decisions guided by its value like a human.Towards this end, we propose to incorporate mental state simulation and value modeling into dialogue agents.First, we build a hybrid mental state parser that extracts information from both the dialogue and event observations and maintains a graphical representation of the agent's mind; Meanwhile, the transformer-based value model learns human preferences from the human value dataset, VALUENET.Empirical results show that the proposed model attains state-of-the-art performance on the dialogue/action/emotion prediction task in the fantasy text-adventure game dataset, LIGHT.We also show example cases to demonstrate: (i) how the proposed mental state parser can assist the agent's decision by grounding on the context like locations and objects, and (ii) how the value model can help the agent make decisions based on its personal priorities. Liang Qiu 0001, Yuan Liang 0001, Pan Lu, Weiyan Shi 0001, Zhou Yu 0005, Song-Chun Zhu |
SIGDIAL | 6 |
| 2022 | Spoken language interaction with robots: Recommendations for future researchabstractWith robotics rapidly advancing, more effective human–robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language interaction capabilities is still very limited. In this article, based on the report of an interdisciplinary workshop convened by the National Science Foundation, we identify key scientific and engineering advances needed to enable effective spoken language interaction with robotics. We make 25 recommendations, involving eight general themes: putting human needs first, better modeling the social and interactive aspects of language, improving robustness, creating new methods for rapid adaptation, better integrating speech and language with other communication modalities, giving speech and language components access to rich representations of the robot’s current knowledge and state, making all components operate in real time, and improving research infrastructure and resources. Research and development that prioritizes these topics will, we believe, provide a solid foundation for the creation of speech-capable robots that are easy and effective for humans to work with. Matthew Marge, Carol Y. Espy-Wilson, Nigel G. Ward, Abeer Alwan, Yoav Artzi, Mohit Bansal, Gilmer L. Blankenship, Joyce Y. Chai, Hal Daumé III, Debadeepta Dey, Mary P. Harper, Thomas Howard, Casey Kennington, Ivana Kruijff-Korbayová, Dinesh Manocha, Cynthia Matuszek, Ross Mead, Raymond J. Mooney, Roger K. Moore, Mari Ostendorf, Heather Pon-Barry, Alexander I. Rudnicky, Matthias Scheutz, Robert St. Amant, Stefanie Tellex, David R. Traum, Zhou Yu 0005 |
Comput. Speech Lang. | 28 |
| 2021 | A Student-Teacher Architecture for Dialog Domain Adaptation Under the Meta-Learning SettingabstractNumerous new dialog domains are being created every day while collecting data for these domains is extremely costly since it involves human interactions. Therefore, it is essential to develop algorithms that can adapt to different domains efficiently when building data-driven dialog models. Most recent research on domain adaption focuses on giving the model a better initialization, rather than optimizing the adaptation process. We propose an efficient domain adaptive task-oriented dialog system model, which incorporates a meta-teacher model to emphasize the different impacts between generated tokens with respect to the context. We first train our base dialog model and meta-teacher model adversarially in a meta-learning setting on rich-resource domains. The meta-teacher learns to quantify the importance of tokens under different contexts across different domains. During adaptation, the meta-teacher guides the dialog model to focus on important tokens in order to achieve better adaptation efficiency. We evaluate our model on two multi-domain datasets, MultiWOZ and Google Schema-Guided Dialogue, and achieve state-of-the-art performance. Kun Qian 0016, Zhou Yu 0005 |
AAAI | 3 |
| 2021 | The R-U-A-Robot Dataset: Helping Avoid Chatbot Deception by Detecting User Questions About Human or Non-Human IdentityabstractDavid Gros, Yu Li, Zhou Yu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. David Gros 0001, Yu Li 0013, Zhou Yu 0005 |
ACL/IJCNLP (1) | 3 |
| 2021 | HERALD: An Annotation Efficient Method to Detect User Disengagement in Social ConversationsabstractWeixin Liang, Kai-Hui Liang, Zhou Yu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Weixin Liang, Kaihui Liang, Zhou Yu 0005 |
ACL/IJCNLP (1) | 3 |
| 2021 | Towards Emotional Support Dialog SystemsabstractSiyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour, Yu Li, Zhou Yu, Yong Jiang, Minlie Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Siyang Liu 0003, Chujie Zheng, Orianna Demasi, Sahand Sabour, Yu Li 0013, Zhou Yu 0005, Yong Jiang 0001, Minlie Huang |
ACL/IJCNLP (1) | 6 |
| 2021 | SocAoG: Incremental Graph Parsing for Social Relation Inference in DialoguesabstractInferring social relations from dialogues is vital for building emotionally intelligent robots to interpret human language better and act accordingly. We model the social network as an And-or Graph, named SocAoG, for the consistency of relations among a group and leveraging attributes as inference cues. Moreover, we formulate a sequential structure prediction task, and propose an $\alpha$-$\beta$-$\gamma$ strategy to incrementally parse SocAoG for the dynamic inference upon any incoming utterance: (i) an $\alpha$ process predicting attributes and relations conditioned on the semantics of dialogues, (ii) a $\beta$ process updating the social relations based on related attributes, and (iii) a $\gamma$ process updating individual's attributes based on interpersonal social relations. Empirical results on DialogRE and MovieGraph show that our model infers social relations more accurately than the state-of-the-art methods. Moreover, the ablation study shows the three processes complement each other, and the case study demonstrates the dynamic relational inference. Liang Qiu 0001, Yuan Liang 0001, Pan Lu, Baolin Peng, Zhou Yu 0005, Ying Nian Wu, Song-Chun Zhu |
ACL/IJCNLP (1) | 6 |
| 2021 | UC2: Universal Cross-Lingual Cross-Modal Vision-and-Language Pre-TrainingabstractVision-and-language pre-training has achieved impressive success in learning multimodal representations between vision and language. To generalize this success to non-English languages, we introduce UC2, the first machine translation-augmented framework for cross-lingual cross-modal representation learning. To tackle the scarcity problem of multilingual captions for image datasets, we first augment existing English-only datasets with other languages via machine translation (MT). Then we extend the standard Masked Language Modeling and Image-Text Matching training objectives to multilingual setting, where alignment between different languages is captured through shared visual context (i.e., using image as pivot). To facilitate the learning of a joint embedding space of images and all languages of interest, we further propose two novel pre-training tasks, namely Masked Region-to-Token Modeling (MRTM) and Visual Translation Language Modeling (VTLM), leveraging MT-enhanced translated data. Evaluation on multilingual image-text retrieval and multilingual visual question answering benchmarks demonstrates that our proposed framework achieves new state of the art on diverse non-English benchmarks while maintaining comparable performance to monolingual pre-trained models on English tasks. Mingyang Zhou 0004, Luowei Zhou, Shuohang Wang, Yu Cheng 0001, Zhou Yu 0005, Jingjing Liu 0001 |
CVPR | 6 |
| 2021 | Alternating Recurrent Dialog Model with Large-scale Pre-trained Language ModelsabstractExisting dialog system models require extensive human annotations and are difficult to generalize to different tasks.The recent success of large pre-trained language models has suggested the effectiveness of incorporating language priors in down-stream NLP tasks.However, how much pre-trained language models can help dialog response generation is still under exploration.In this paper, we propose a simple, general, and effective framework: Alternating Recurrent Dialog Model (ARDM) 1 .ARDM models each speaker separately and takes advantage of large pre-trained language models.It requires no supervision from human annotations such as belief states or dialog acts to achieve effective conversations.ARDM outperforms or is on par with the state-of-theart methods on two popular task-oriented dialog datasets: CamRest676 and MultiWOZ.Moreover, we can generalize ARDM to more challenging, non-collaborative tasks such as persuasion.In the PersuasionForGood task, ARDM is capable of generating human-like responses to persuade people to donate to a charity. Qingyang Wu, Yichi Zhang 0001, Yu Li 0013, Zhou Yu 0005 |
EACL | 4 |
| 2021 | MIDAS: A Dialog Act Annotation Scheme for Open Domain HumanMachine Spoken ConversationsabstractDialog act prediction in open-domain conversations is an essential language comprehension task for both dialog system building and discourse analysis.Previous dialog act schemes, such as SWBD-DAMSL, are designed mainly for discourse analysis in humanhuman conversations.In this paper, we present a dialog act annotation scheme, MIDAS (Machine Interaction Dialog Act Scheme), targeted at open-domain human-machine conversations.MIDAS is designed to assist machines to improve their ability to understand human partners.MIDAS has a hierarchical structure and supports multi-label annotations.We collected and annotated a large open-domain human-machine spoken conversation dataset (consisting of 24K utterances).To validate our scheme, we leveraged transfer learning methods to train a multi-label dialog act prediction model and reached an F1 score of 0.79. 1 Dian Yu 0002, Zhou Yu 0005 |
EACL | 2 |
| 2021 | Zero-Shot Dialogue State Tracking via Cross-Task TransferabstractZhaojiang Lin, Bing Liu, Andrea Madotto, Seungwhan Moon, Zhenpeng Zhou, Paul Crook, Zhiguang Wang, Zhou Yu, Eunjoon Cho, Rajen Subba, Pascale Fung. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhaojiang Lin, Andrea Madotto, Seungwhan Moon, Zhenpeng Zhou, Paul A. Crook, Zhiguang Wang, Zhou Yu 0005, Eunjoon Cho, Rajen Subba, Pascale Fung |
EMNLP (1) | 8 |
| 2021 | Continual Learning in Task-Oriented Dialogue SystemsabstractAndrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul Crook, Bing Liu, Zhou Yu, Eunjoon Cho, Pascale Fung, Zhiguang Wang. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Andrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul A. Crook, Zhou Yu 0005, Eunjoon Cho, Pascale Fung, Zhiguang Wang |
EMNLP (1) | 7 |
| 2021 | Evaluation of In-Person Counseling Strategies To Develop Physical Activity Chatbot for WomenabstractArtificial intelligence chatbots are the vanguard in technology-based intervention to change people's behavior.To develop intervention chatbots, the first step is to understand natural language conversation strategies in human conversation.This work introduces an intervention conversation dataset collected from a real-world physical activity intervention program for women.We designed comprehensive annotation schemes in four dimensions (domain, strategy, social exchange, and taskfocused exchange) and annotated a subset of dialogs.We built a strategy classifier with context information to detect strategies from both trainers and participants based on the annotation.To understand how human intervention induces effective behavior changes, we analyzed the relationships between the intervention strategies and the participants' changes in the barrier and social support for physical activity.We also analyzed how participant's baseline weight correlates to the amount of occurrence of the corresponding strategy.This work lays the foundation for developing a personalized physical activity intervention bot. 1 Kaihui Liang, Patrick L. Lange, Yoshimi Fukuoka, Zhou Yu 0005 |
SIGDIAL | 6 |
| 2021 | Annotation Inconsistency and Entity Bias in MultiWOZabstractMultiWOZ (Budzianowski et al., 2018) is one of the most popular multi-domain taskoriented dialog datasets, containing 10K+ annotated dialogs covering eight domains.It has been widely accepted as a benchmark for various dialog tasks, e.g., dialog state tracking (DST), natural language generation (NLG) and end-to-end (E2E) dialog modeling.In this work, we identify an overlooked issue with dialog state annotation inconsistencies in the dataset, where a slot type is tagged inconsistently across similar dialogs leading to confusion for DST modeling.We propose an automated correction for this issue, which is present in 70% of the dialogs.Additionally, we notice that there is significant entity bias in the dataset (e.g., "cambridge" appears in 50% of the destination cities in the train domain).The entity bias can potentially lead to named entity memorization in generative models, which may go unnoticed as the test set suffers from a similar entity bias as well.We release a new test set with all entities replaced with unseen entities.Finally, we benchmark joint goal accuracy (JGA) of the state-of-theart DST baselines on these modified versions of the data.Our experiments show that the annotation inconsistency corrections lead to 7-10% improvement in JGA.On the other hand, we observe a 29% drop in JGA when models are evaluated on the new test set with unseen entities.* The work of KQ and ZY was done as a research intern and a visiting research scientist at Facebook AI. Kun Qian 0016, Ahmad Beirami, Zhouhan Lin, Ankita De, Alborz Geramifard, Zhou Yu 0005, Chinnadhurai Sankar |
SIGDIAL | 6 |
| 2020 | End-to-End Trainable Non-Collaborative Dialog SystemabstractEnd-to-end task-oriented dialog models have achieved promising performance on collaborative tasks where users willingly coordinate with the system to complete a given task. While in non-collaborative settings, for example, negotiation and persuasion, users and systems do not share a common goal. As a result, compared to collaborate tasks, people use social content to build rapport and trust in these non-collaborative settings in order to advance their goals. To handle social content, we introduce a hierarchical intent annotation scheme, which can be generalized to different non-collaborative dialog tasks. Building upon TransferTransfo (Wolf et al. 2019), we propose an end-to-end neural network model to generate diverse coherent responses. Our model utilizes intent and semantic slots as the intermediate sentence representation to guide the generation process. In addition, we design a filter to select appropriate responses based on whether these intermediate representations fit the designed task and conversation constraints. Our non-collaborative dialog model guides users to complete the task while simultaneously keeps them engaged. We test our approach on our newly proposed AntiScam dataset and an existing PersuasionForGood dataset. Both automatic and human evaluations suggest that our model outperforms multiple baselines in these two non-collaborative tasks. Yu Li 0013, Kun Qian 0016, Weiyan Shi 0001, Zhou Yu 0005 |
AAAI | 4 |
| 2020 | MOSS: End-to-End Dialog System Framework with Modular SupervisionabstractA major bottleneck in training end-to-end task-oriented dialog system is the lack of data. To utilize limited training data more efficiently, we propose Modular Supervision Network (MOSS), an encoder-decoder training framework that could incorporate supervision from various intermediate dialog system modules including natural language understanding, dialog state tracking, dialog policy learning and natural language generation. With only 60% of the training data, MOSS-all (i.e., MOSS with supervision from all four dialog modules) outperforms state-of-the-art models on CamRest676. Moreover, introducing modular supervision has even bigger benefits when the dialog task has a more complex dialog state and action space. With only 40% of the training data, MOSS-all outperforms the state-of-the-art model on a complex laptop network trouble shooting dataset, LaptopNetwork, that we introduced. LaptopNetwork consists of conversations between real customers and customer service agents in Chinese. Moreover, MOSS framework can accommodate dialogs that have supervision from different dialog modules at both framework level and model level. Therefore, MOSS is extremely flexible to update in real-world deployment. Weixin Liang, Youzhi Tian, Chengcai Chen, Zhou Yu 0005 |
AAAI | 4 |
| 2020 | Filling Conversation Ellipsis for Better Social Dialog UnderstandingabstractThe phenomenon of ellipsis is prevalent in social conversations. Ellipsis increases the difficulty of a series of downstream language understanding tasks, such as dialog act prediction and semantic role labeling. We propose to resolve ellipsis through automatic sentence completion to improve language understanding. However, automatic ellipsis completion can result in output which does not accurately reflect user intent. To address this issue, we propose a method which considers both the original utterance that has ellipsis and the automatically completed utterance in dialog act and semantic role labeling tasks. Specifically, we first complete user utterances to resolve ellipsis using an end-to-end pointer network model. We then train a prediction model using both utterances containing ellipsis and our automatically completed utterances. Finally, we combine the prediction results from these two utterances using a selection model that is guided by expert knowledge. Our approach improves dialog act prediction and semantic role labeling by 1.3% and 2.5% in F1 score respectively in social conversations. We also present an open-domain human-machine conversation dataset with manually completed user utterances and annotated semantic role labeling after manual completion. Xiyuan Zhang 0001, Chengxi Li 0002, Dian Yu 0002, Sam Davidson, Zhou Yu 0005 |
AAAI | 5 |
| 2020 | Task-Oriented Dialog Systems That Consider Multiple Appropriate Responses under the Same ContextabstractConversations have an intrinsic one-to-many property, which means that multiple responses can be appropriate for the same dialog context. In task-oriented dialogs, this property leads to different valid dialog policies towards task completion. However, none of the existing task-oriented dialog generation approaches takes this property into account. We propose a Multi-Action Data Augmentation (MADA) framework to utilize the one-to-many property to generate diverse appropriate dialog responses. Specifically, we first use dialog states to summarize the dialog history, and then discover all possible mappings from every dialog state to its different valid system actions. During dialog system training, we enable the current dialog state to map to all valid system actions discovered in the previous process to create additional state-action pairs. By incorporating these additional pairs, the dialog policy learns a balanced action distribution, which further guides the dialog model to generate diverse responses. Experimental results show that the proposed framework consistently improves dialog policy diversity, and results in improved response diversity and appropriateness. Our model obtains state-of-the-art results on MultiWOZ. Yichi Zhang 0001, Zhijian Ou, Zhou Yu 0005 |
AAAI | 3 |
| 2020 | Paraphrase Augmented Task-Oriented Dialog GenerationabstractNeural generative models have achieved promising performance on dialog generation tasks if given a huge data set.However, the lack of high-quality dialog data and the expensive data annotation process greatly limit their application in real-world settings.We propose a paraphrase augmented response generation (PARG) framework that jointly trains a paraphrase model and a response generation model to improve the dialog generation performance.We also design a method to automatically construct paraphrase training data set based on dialog state and dialog act labels.PARG is applicable to various dialog generation models, such as TSCP (Lei et al., 2018) and DAMD (Zhang et al., 2019).Experimental results show that the proposed framework improves these state-of-the-art dialog models further on CamRest676 and MultiWOZ.PARG also significantly outperforms other data augmentation methods in dialog generation tasks, especially under low resource settings.1 2 Silin Gao, Yichi Zhang 0001, Zhijian Ou, Zhou Yu 0005 |
ACL | 4 |
| 2020 | SAS: Dialogue State Tracking via Slot Attention and Slot Information SharingabstractDialogue state tracker is responsible for inferring user intentions through dialogue history.Previous methods have difficulties in handling dialogues with long interaction context, due to the excessive information.We propose a Dialogue State Tracker with Slot Attention and Slot Information Sharing (SAS) to reduce redundant information's interference and improve long dialogue context tracking.Specially, we first apply a Slot Attention to learn a set of slot-specific features from the original dialogue and then integrate them using a Slot Information Sharing.The sharing improve the models ability to deduce value from related slots.Our model yields a significantly improved performance compared to previous state-of-the-art models on the Multi-WOZ dataset. Jiaying Hu, Yan Yang 0008, Chencai Chen, Liang He 0001, Zhou Yu 0005 |
ACL | 5 |
| 2020 | Effects of Persuasive Dialogues: Testing Bot Identities and Inquiry StrategiesabstractIntelligent conversational agents, or chatbots, can take on various identities and are increasingly engaging in more human-centered conversations with persuasive goals. However, little is known about how identities and inquiry strategies influence the conversation's effectiveness. We conducted an online study involving 790 participants to be persuaded by a chatbot for charity donation. We designed a two by four factorial experiment (two chatbot identities and four inquiry strategies) where participants were randomly assigned to different conditions. Findings showed that the perceived identity of the chatbot had significant effects on the persuasion outcome (i.e., donation) and interpersonal perceptions (i.e., competence, confidence, warmth, and sincerity). Further, we identified interaction effects among perceived identities and inquiry strategies. We discuss the findings for theoretical and practical implications for developing ethical and effective persuasive chatbots. Our published data, codes, and analyses serve as the first step towards building competent ethical persuasive chatbots. Weiyan Shi 0001, Saurav Sahay, Zhou Yu 0005 |
CHI | 6 |
| 2020 | INSPIRED: Toward Sociable Recommendation Dialog SystemsabstractIn recommendation dialogs, humans commonly disclose their preference and make recommendations in a friendly manner.However, this is a challenge in developing a sociable recommendation dialog system, due to the lack of dialog dataset annotated with such sociable strategies.Therefore, we present INSPIRED, a new dataset of 1,001 human-human dialogs for movie recommendation with measures for successful recommendations.To better understand how humans make recommendations in communication, we design an annotation scheme related to recommendation strategies based on social science theories and annotate these dialogs.Our analysis shows that sociable recommendation strategies, such as sharing personal opinions or communicating with encouragement, more frequently lead to successful recommendations.Based on our dataset, we train end-to-end recommendation dialog systems with and without our strategy labels.In both automatic and human evaluation, our model with strategy incorporation outperforms the baseline model.This work is a first step for building sociable recommendation dialog systems with a basis of social science theories 1 . Shirley Anugrah Hayati, Dongyeop Kang, Qingxiaoyang Zhu, Weiyan Shi 0001, Zhou Yu 0005 |
EMNLP (1) | 5 |
| 2020 | Structured Attention for Unsupervised Dialogue Structure InductionabstractInducing a meaningful structural representation from one or a set of dialogues is a crucial but challenging task in computational linguistics. Advancement made in this area is critical for dialogue system design and discourse analysis. It can also be extended to solve grammatical inference. In this work, we propose to incorporate structured attention layers into a Variational Recurrent Neural Network (VRNN) model with discrete latent states to learn dialogue structure in an unsupervised fashion. Compared to a vanilla VRNN, structured attention enables a model to focus on different parts of the source sentence embeddings while enforcing a structural inductive bias. Experiments show that on two-party dialogue datasets, VRNN with structured attention learns semantic structures that are similar to templates used to generate this dialogue corpus. While on multi-party dialogue datasets, our model learns an interactive structure demonstrating its capability of distinguishing speakers or addresses, automatically disentangling dialogues without explicit human annotation. Liang Qiu 0001, Weiyan Shi 0001, Yuan Liang 0001, Feng Shi 0006, Zhou Yu 0005, Song-Chun Zhu |
EMNLP (1) | 7 |
| 2020 | Augmenting Non-Collaborative Dialog Systems with Explicit Semantic and Strategic Dialog History
Yiheng Zhou, Yulia Tsvetkov, Alan W. Black, Zhou Yu 0005 |
ICLR | 4 |
| 2020 | Code to Comment "Translation": Data, Metrics, Baselining & EvaluationabstractThe relationship of comments to code, and in particular, the task of generating useful comments given the code, has long been of interest. The earliest approaches have been based on strong syntactic theories of comment-structures, and relied on textual templates. More recently, researchers have applied deep-learning methods to this task---specifically, trainable generative translation models which are known to work very well for Natural Language translation (e.g., from German to English). We carefully examine the underlying assumption here: that the task of generating comments sufficiently resembles the task of translating between natural languages, and so similar models and evaluation metrics could be used. We analyze several recent code-comment datasets for this task: CodeNN, DeepCom, FunCom, and DocString. We compare them with WMT19, a standard dataset frequently used to train state-of-the-art natural language translators. We found some interesting differences between the code-comment data and the WMT19 natural language data. Next, we describe and conduct some studies to calibrate BLEU (which is commonly used as a measure of comment quality). using "affinity pairs" of methods, from different projects, in the same project, in the same class, etc; Our study suggests that the current performance on some datasets might need to be improved substantially. We also argue that fairly naive information retrieval (IR) methods do well enough at this task to be considered a reasonable baseline. Finally, we make some suggestions on how our findings might be used in future research in this area. David Gros 0001, Hariharan Sezhiyan, Premkumar T. Devanbu, Zhou Yu 0005 |
ASE | 4 |
| 2020 | Generating Emotional Social Chatbot Responses with a Consistent Speaking Style
Jun Zhang 0098, Yan Yang 0008, Chengcai Chen, Liang He 0001, Zhou Yu 0005 |
NLPCC (2) | 5 |
| 2019 | Domain Adaptive Dialog Generation via Meta LearningabstractDomain adaptation is an essential task in dialog system building because there are so many new dialog tasks created for different needs every day.Collecting and annotating training data for these new tasks is costly since it involves real user interactions.We propose a domain adaptive dialog generation method based on meta-learning (DAML).DAML is an end-to-end trainable dialog system model that learns from multiple rich-resource tasks and then adapts to new domains with minimal training samples.We train a dialog system model using multiple rich-resource singledomain dialog data by applying the modelagnostic meta-learning algorithm to dialog domain.The model is capable of learning a competitive dialog system on a new domain with only a few training examples in an efficient manner.The two-step gradient updates in DAML enable the model to learn general features across multiple tasks.We evaluate our method on a simulated dialog dataset and achieve state-of-the-art performance, which is generalizable to new tasks. Kun Qian 0016, Zhou Yu 0005 |
ACL (1) | 2 |
| 2019 | Persuasion for Good: Towards a Personalized Persuasive Dialogue System for Social GoodabstractDeveloping intelligent persuasive conversational agents to change people's opinions and actions for social good is the frontier in advancing the ethical development of automated dialogue systems.To do so, the first step is to understand the intricate organization of strategic disclosures and appeals employed in human persuasion conversations.We designed an online persuasion task where one participant was asked to persuade the other to donate to a specific charity.We collected a large dataset with 1,017 dialogues and annotated emerging persuasion strategies from a subset.Based on the annotation, we built a baseline classifier with context information and sentence-level features to predict the 10 persuasion strategies used in the corpus.Furthermore, to develop an understanding of personalized persuasion processes, we analyzed the relationships between individuals' demographic and psychological backgrounds including personality, morality, value systems, and their willingness for donation.Then, we analyzed which types of persuasion strategies led to a greater amount of donation depending on the individuals' personal backgrounds.This work lays the ground for developing a personalized persuasive dialogue system. 1 Weiyan Shi 0001, Richard Kim, Zhou Yu 0005 |
ACL (1) | 7 |
| 2019 | Dependency Parsing for Spoken Dialog SystemsabstractSam Davidson, Dian Yu, Zhou Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sam Davidson, Dian Yu 0002, Zhou Yu 0005 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | How to Build User Simulators to Train RL-based Dialog SystemsabstractWeiyan Shi, Kun Qian, Xuewei Wang, Zhou Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Weiyan Shi 0001, Kun Qian 0016, Zhou Yu 0005 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Building Task-Oriented Visual Dialog Systems Through Alternative Optimization Between Dialog Policy and Language GenerationabstractMingyang Zhou, Josh Arnold, Zhou Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mingyang Zhou 0004, Josh Arnold, Zhou Yu 0005 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Dependent Multilevel Interaction Network for Natural Language Inference
Yan Yang 0008, Qinmin Hu, Chengcai Chen, Liang He 0001, Zhou Yu 0005 |
ICANN (4) | 7 |
| 2019 | Knowledge Adaptive Neural Network for Natural Language InferenceabstractNatural language inference (NLI) has received widespread attention in recent years due to its contribution to various natural language processing tasks, such as question answering, abstract text summarization, and video caption. Most existing works focus on modeling the sentence interaction information, while the use of commonsense knowledge is not well studied for NLI. In this paper, we propose knowledge adaptive neural network (KANN) that adaptively incorporates commonsense knowledge at sentence encoding and inference stages. We first perform knowledge collection and representation to identify the relevant knowledge. Then we use a knowledge absorption gate to embed knowledge into neural network models. Experiments on two benchmark datasets, namely SNLI and MultiNLI for natural language inference, show the advantages of our proposed model. Furthermore, our model is comparable to if not better than the recent neural network based approaches on NLI. Qi Zhang 0001, Yan Yang 0008, Chengcai Chen, Liang He 0001, Zhou Yu 0005 |
IJCNN | 5 |
| 2019 | A Large-Scale User Study of an Alexa Prize Chatbot: Effect of TTS Dynamism on Perceived Quality of Social DialogabstractThis study tests the effect of cognitiveemotional expression in an Alexa text-tospeech (TTS) voice on users' experience with a social dialog system.We systematically introduced emotionally expressive interjections (e.g., "Wow!") and filler words (e.g., "um", "mhmm") in an Amazon Alexa Prize socialbot, Gunrock.We tested whether these TTS manipulations improved users' ratings of their conversation across thousands of real user interactions (n=5,527).Results showed that interjections and fillers each improved users' holistic ratings, an improvement that further increased if the system used both manipulations.A separate perception experiment corroborated the findings from the user study, with improved social ratings for conversations including interjections; however, no positive effect was observed for fillers, suggesting that the role of the rater in the conversation-as active participant or external listener-is an important factor in assessing social dialogs. Michelle Cohn, Zhou Yu 0005 |
SIGdial | 3 |
| 2018 | Sentiment Adaptive End-to-End Dialog SystemsabstractEnd-to-end learning framework is useful for building dialog systems for its simplicity in training and efficiency in model updating.However, current end-to-end approaches only consider user semantic inputs in learning and under-utilize other user information.Therefore, we propose to include user sentiment obtained through multimodal information (acoustic, dialogic and textual), in the end-to-end learning framework to make systems more user-adaptive and effective.We incorporated user sentiment information in both supervised and reinforcement learning settings.In both settings, adding sentiment information reduced the dialog length and improved the task success rate on a bus information search task.This work is the first attempt to incorporate multimodal user information in the adaptive end-toend dialog system training framework and attained state-of-the-art performance. Weiyan Shi 0001, Zhou Yu 0005 |
ACL (1) | 2 |
| 2018 | A Visual Attention Grounding Neural Model for Multimodal Machine TranslationabstractWe introduce a novel multimodal machine translation model that utilizes parallel visual and textual information.Our model jointly optimizes the learning of a shared visuallanguage embedding and a translator.The model leverages a visual attention grounding mechanism that links the visual semantics with the corresponding textual semantics.Our approach achieves competitive state-of-the-art results on the Multi30K and the Ambiguous COCO datasets.We also collected a new multilingual multimodal product description dataset to simulate a real-world international online shopping scenario.On this dataset, our visual attention grounding model outperforms other methods by a large margin. Mingyang Zhou 0004, Runxiang Cheng, Yong Jae Lee, Zhou Yu 0005 |
EMNLP | 4 |
| 2017 | Learning Conversational Systems that Interleave Task and Non-Task ContentabstractTask-oriented dialog systems have been applied in various tasks, such as automated personal assistants, customer service providers and tutors. These systems work well when users have clear and explicit intentions that are well-aligned to the systems' capabilities. However, they fail if users intentions are not explicit.To address this shortcoming, we propose a framework to interleave non-task content (i.e.everyday social conversation) into task conversations. When the task content fails, the system can still keep the user engaged with the non-task content. We trained a policy using reinforcement learning algorithms to promote long-turn conversation coherence and consistency, so that the system can have smooth transitions between task and non-task content.To test the effectiveness of the proposed framework, we developed a movie promotion dialog system. Experiments with human users indicate that a system that interleaves social and task content achieves a better task success rate and is also rated as more engaging compared to a pure task-oriented system. Zhou Yu 0005, Alexander I. Rudnicky, Alan W. Black |
IJCAI | 1 |
| 2016 | User Engagement Study with Virtual Agents Under Different Cultural Contexts
Zhou Yu 0005, Xinrui He, Alan W. Black, Alexander I. Rudnicky |
IVA | 1 |
| 2016 | A Wizard-of-Oz Study on A Non-Task-Oriented Dialog Systems That Reacts to User EngagementabstractIn this paper, we describe a system that reacts to both possible system breakdowns and low user engagement with a set of conversational strategies.These general strategies reduce the number of inappropriate responses and produce better user engagement.We also found that a system that reacts to both possible system breakdowns and low user engagement is rated by both experts and non-experts as having better overall user engagement compared to a system that only reacts to possible system breakdowns.We argue that for non-task-oriented systems we should optimize on both system response appropriateness and user engagement.We also found that apart from making the system response appropriate, funny and provocative responses can also lead to better user engagement.On the other hand, short appropriate responses, such as "Yes" or "No" can lead to decreased user engagement.We will use these findings to further improve our system. Zhou Yu 0005, Leah Nicolich-Henkin, Alan W. Black, Alexander I. Rudnicky |
SIGDIAL Conference | 1 |
| 2016 | Strategy and Policy Learning for Non-Task-Oriented Conversational SystemsabstractWe propose a set of generic conversational strategies to handle possible system breakdowns in non-task-oriented dialog systems.We also design policies to select these strategies according to dialog context.We combine expert knowledge and the statistical findings derived from data in designing these policies.The policy learned via reinforcement learning outperforms the random selection policy and the locally greedy policy in both simulated and real-world settings.In addition, we propose three metrics for conversation quality evaluation which consider both the local and global quality of the conversation. Zhou Yu 0005, Ziyu Xu 0001, Alan W. Black, Alexander I. Rudnicky |
SIGDIAL Conference | 1 |
| 2015 | Using bidirectional lstm recurrent neural networks to learn high-level abstractions of sequential features for automated scoring of non-native spontaneous speechabstractWe introduce a new method to grade non-native spoken language tests automatically. Traditional automated response grading approaches use manually engineered time-aggregated features (such as mean length of pauses). We propose to incorporate general time-sequence features (such as pitch) which preserve more information than time-aggregated features and do not require human effort to design. We use a type of recurrent neural network to jointly optimize the learning of high level abstractions from time-sequence features with the time-aggregated features. We first automatically learn high level abstractions from time-sequence features with a Bidirectional Long Short Term Memory (BLSTM) and then combine the high level abstractions with time-aggregated features in a Multilayer Perceptron (MLP)/Linear Regression (LR). We optimize the BLSTM and the MLP/LR jointly. We find such models reach the best performance in terms of correlation with human raters. We also find that when there are limited time-aggregated features available, our model that incorporates time-sequence features improves performance drastically. Zhou Yu 0005, Vikram Ramanarayanan, David Suendermann-Oeft, Klaus Zechner, Lei Chen 0004, Jidong Tao, Aliaksei Ivanou, Yao Qian |
ASRU | 1 |
| 2013 | Automatic Prediction of Friendship via Multi-model Dyadic Features
Zhou Yu 0005, David Gerritsen, Amy Ogan, Alan W. Black, Justine Cassell |
SIGDIAL Conference | 1 |