Keqing He 0001

dblp:79/2314-1 · DBLP profile ↗
← Back
44ranked-venue papers
8as first author
37since 2021 · last 2026
0000-0002-8831-560XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 38 · 8 first-author · 31 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SelfCorrect-Agent: Toward robust and generalizable LLM-based agents
Keqing He 0001, Dayuan Fu, Lele Yang, Weiran Xu
Neurocomputing1
2025 PreAct: Prediction Enhances Agent's Planning Ability
abstract
Addressing the disparity between predictions and actual results can enable individuals to expand their thought processes and stimulate self-reflection, thus promoting accurate planning. In this research, we present PreAct, an agent framework that integrates prediction, reasoning, and action. By utilizing the information derived from predictions, the large language model (LLM) agent can provide a wider range and more strategically focused reasoning. This leads to more efficient actions that aid the agent in accomplishing intricate tasks. Our experimental results show that PreAct surpasses the ReAct method in completing complex tasks and that PreAct’s performance can be further improved when paired with other memory or selection strategy techniques. We presented the model with varying quantities of historical predictions and discovered that these predictions consistently enhance LLM planning. The variances in single-step reasoning between PreAct and ReAct indicate that PreAct indeed has benefits in terms of diversity and strategic orientation over ReAct.
Dayuan Fu, Jianzhao Huang, Guanting Dong 0001, Yejie Wang, Keqing He 0001, Weiran Xu
COLING6
2025 AgentRefine: Enhancing Agent Generalization through Refinement Tuning
abstract
Large Language Model (LLM) based agents have proved their ability to perform complex tasks like humans. However, there is still a large gap between open-sourced LLMs and commercial models like the GPT series. In this paper, we focus on improving the agent generalization capabilities of LLMs via instruction tuning. We first observe that the existing agent training corpus exhibits satisfactory results on held-in evaluation sets but fails to generalize to held-out sets. These agent-tuning works face severe formatting errors and are frequently stuck in the same mistake for a long while. We analyze that the poor generalization ability comes from overfitting to several manual agent environments and a lack of adaptation to new situations. They struggle with the wrong action steps and can not learn from the experience but just memorize existing observation-action relations. Inspired by the insight, we propose a novel AgentRefine framework for agent-tuning. The core idea is to enable the model to learn to correct its mistakes via observation in the trajectory. Specifically, we propose an agent synthesis framework to encompass a diverse array of environments and tasks and prompt a strong LLM to refine its error action according to the environment feedback. AgentRefine significantly outperforms state-of-the-art agent-tuning work in terms of generalization ability on diverse agent tasks. It also has better robustness facing perturbation and can generate diversified thought in inference. Our findings establish the correlation between agent generalization and self-refinement and provide a new paradigm for future research.
Dayuan Fu, Keqing He 0001, Yejie Wang, Wentao Hong, Zhuoma Gongque, Weihao Zeng 0003, Jingang Wang, Weiran Xu
ICLR2
2025 Assessing and Post-Processing Black Box Large Language Models for Knowledge Editing
abstract
The rapid evolution of the Web as a key platform for information dissemination has led to the growing integration of large language models (LLMs) in Web-based applications. However, the swift changes in web content present challenges in maintaining these models' relevance and accuracy. The task of Knowledge Editing (KE) is aimed at efficiently and precisely adjusting the behavior of large language models (LLMs) to update specific knowledge while minimizing any adverse effects on other knowledge. Current research predominantly concentrates on editing white-box LLMs, neglecting a significant scenario: editing black-box LLMs, where access is limited to interfaces and only textual output is provided. In this paper, we initially officially introduce KE on black-box LLMs, followed by presenting a thorough evaluation framework. This framework operates without requiring logits and considers pre- and post-edit consistency, addressing the limitations of current evaluations that are inadequate for black-box LLMs editing and lack comprehensiveness. To address privacy leaks of editing data and style over-editing in existing approaches, we propose a new postEdit framework. postEdit incorporates a retrieval mechanism for editing knowledge and a purpose-trained editing plugin called post-editor, ensuring privacy through downstream processing and maintaining textual style consistency via fine-grained editing. Experiments and analysis conducted on two benchmarks show that postEdit surpasses all baselines and exhibits robust generalization, notably enhancing style retention by an average of +20.82%. Our code is available on github https://github.com/songxiaoshuai/postEdit.
Xiaoshuai Song, Keqing He 0001, Guanting Dong 0001, Yutao Mou, Jinxu Zhao, Weiran Xu
WWW3
2025 Towards robust and generalizable training: An empirical study of noisy slot filling for input perturbations
Guanting Dong 0001, Xiaoshuai Song, Jiachi Liu, Liwen Wang 0007, Zechen Wang, Shanglin Lei, Keqing He 0001, Weiran Xu
Neurocomputing8
2025 FuseMind: Fusing reflection and prediction elevates agent's reasoning capabilities
Xiufa Ma, Xinxin Ge, Heyang Xu, Dayuan Fu, Zhexu Wang, Keqing He 0001, Weiran Xu
Neurocomputing7
2024 DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning
abstract
Yejie Wang, Keqing He, Guanting Dong, Pei Wang, Weihao Zeng, Muxi Diao, Weiran Xu, Jingang Wang, Mengdi Zhang, Xunliang Cai. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yejie Wang, Keqing He 0001, Guanting Dong 0001, Weihao Zeng 0003, Muxi Diao, Weiran Xu, Jingang Wang, Mengdi Zhang 0002
ACL (1)2
2024 Beyond the Known: Investigating LLMs Performance on Out-of-Domain Intent Detection
abstract
Out-of-domain (OOD) intent detection aims to examine whether the user’s query falls outside the predefined domain of the system, which is crucial for the proper functioning of task-oriented dialogue (TOD) systems. Previous methods address it by fine-tuning discriminative models. Recently, some studies have been exploring the application of large language models (LLMs) represented by ChatGPT to various downstream tasks, but it is still unclear for their ability on OOD detection task.This paper conducts a comprehensive evaluation of LLMs under various experimental settings, and then outline the strengths and weaknesses of LLMs. We find that LLMs exhibit strong zero-shot and few-shot capabilities, but is still at a disadvantage compared to models fine-tuned with full resource. More deeply, through a series of additional analysis experiments, we discuss and summarize the challenges faced by LLMs and provide guidance for future work including injecting domain knowledge, strengthening knowledge transfer from IND(In-domain) to OOD, and understanding long instructions.
Keqing He 0001, Yejie Wang, Xiaoshuai Song, Yutao Mou, Jingang Wang, Yunsen Xian, Weiran Xu
LREC/COLING2
2024 BootTOD: Bootstrap Task-oriented Dialogue Representations by Aligning Diverse Responses
abstract
Pre-trained language models have been successful in many scenarios. However, their usefulness in task-oriented dialogues is limited due to the intrinsic linguistic differences between general text and task-oriented dialogues. Current task-oriented dialogue pre-training methods rely on a contrastive framework, which faces challenges such as selecting true positives and hard negatives, as well as lacking diversity. In this paper, we propose a novel dialogue pre-training model called BootTOD. It learns task-oriented dialogue representations via a self-bootstrapping framework. Unlike contrastive counterparts, BootTOD aligns context and context+response representations and dismisses the requirements of contrastive pairs. BootTOD also uses multiple appropriate response targets to model the intrinsic one-to-many diversity of human conversations. Experimental results show that BootTOD outperforms strong TOD baselines on diverse downstream dialogue tasks.
Weihao Zeng 0003, Keqing He 0001, Yejie Wang, Dayuan Fu, Weiran Xu
LREC/COLING2
2024 How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good Data
abstract
Yejie Wang, Keqing He, Dayuan Fu, Zhuoma GongQue, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong, Muxi Diao, Jingang Wang, Mengdi Zhang, Xunliang Cai, Weiran Xu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yejie Wang, Keqing He 0001, Dayuan Fu, Zhuoma Gongque, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong 0001, Muxi Diao, Jingang Wang, Mengdi Zhang 0002, Weiran Xu
EMNLP2
2024 Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models
abstract
The scaling of large language models (LLMs) is a critical research area for the efficiency and effectiveness of model training and deployment.Our work investigates the transferability and discrepancies of scaling laws between Dense Models and Mixture of Experts (MoE) models.Through a combination of theoretical analysis and extensive experiments, including consistent loss scaling, optimal batch size and learning rate scaling, and resource allocation strategies scaling, our findings reveal that the power-law scaling framework also applies to MoE Models, indicating that the fundamental principles governing the scaling behavior of these models are preserved, even though the architecture differs.Additionally, MoE Models demonstrate superior generalization, resulting in lower testing losses with the same training compute budget compared to Dense Models.These findings indicate the scaling consistency and transfer generalization capabilities of MoE Models, providing new insights for optimizing MoE Model training and deployment strategies.
Zhengyu Chen 0001, Keqing He 0001, Min Zhang 0068, Jingang Wang
EMNLP4
2024 What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
abstract
Instruction tuning is a standard technique employed to align large language models to end tasks and user preferences after the initial pretraining phase. Recent research indicates the critical role of data engineering in instruction tuning -- when appropriately selected, only limited data is necessary to achieve superior performance. However, we still lack a principled understanding of what makes good instruction tuning data for alignment, and how we should select data automatically and effectively. In this work, we delve deeply into automatic data selection strategies for alignment. We start with controlled studies to measure data across three dimensions: complexity, quality, and diversity, along which we examine existing methods and introduce novel techniques for enhanced data measurement. Subsequently, we propose a simple strategy to select data samples based on the measurement. We present Deita (short for Data-Efficient Instruction Tuning for Alignment), a series of models fine-tuned from LLaMA models using data samples automatically selected with our proposed approach. When assessed through both automatic metrics and human evaluation, Deita performs better or on par with the state-of-the-art open-source alignment models such as Vicuna and WizardLM with only 6K training data samples -- 10x less than the data used in the baselines. We anticipate this work to provide clear guidelines and tools on automatic data selection, aiding researchers and practitioners in achieving data-efficient alignment.
Wei Liu 0131, Weihao Zeng 0003, Keqing He 0001, Yong Jiang 0001, Junxian He
ICLR3
2023 Decoupling Pseudo Label Disambiguation and Representation Learning for Generalized Intent Discovery
abstract
Yutao Mou, Xiaoshuai Song, Keqing He, Chen Zeng, Pei Wang, Jingang Wang, Yunsen Xian, Weiran Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yutao Mou, Xiaoshuai Song, Keqing He 0001, Jingang Wang, Yunsen Xian, Weiran Xu
ACL (1)3
2023 FutureTOD: Teaching Future Knowledge to Pre-trained Language Model for Task-Oriented Dialogue
abstract
Weihao Zeng, Keqing He, Yejie Wang, Chen Zeng, Jingang Wang, Yunsen Xian, Weiran Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Weihao Zeng 0003, Keqing He 0001, Yejie Wang, Jingang Wang, Yunsen Xian, Weiran Xu
ACL (1)2
2023 Seen to Unseen: Exploring Compositional Generalization of Multi-Attribute Controllable Dialogue Generation
abstract
Existing controllable dialogue generation work focuses on the single-attribute control and lacks generalization capability to out-of-distribution multiple attribute combinations.In this paper, we explore the compositional generalization for multi-attribute controllable dialogue generation where a model can learn from seen attribute values and generalize to unseen combinations.We propose a prompt-based disentangled controllable dialogue generation model, DCG.It learns attribute concept composition by generating attribute-oriented prompt vectors and uses a disentanglement loss to disentangle different attributes for better generalization.Besides, we design a unified reference-free evaluation framework for multiple attributes with different levels of granularities.Experiment results on two benchmarks prove the effectiveness of our method and the evaluation metric.* The first two authors contribute equally.Weiran Xu is the corresponding author.
Weihao Zeng 0003, Keqing He 0001, Ruotong Geng, Jingang Wang, Wei Wu 0014, Weiran Xu
ACL (1)3
2023 A Multi-Task Semantic Decomposition Framework with Task-specific Pre-training for Few-Shot NER
abstract
The objective of few-shot named entity recognition is to identify named entities with limited labeled instances. Previous works have primarily focused on optimizing the traditional token-wise classification framework, while neglecting the exploration of information based on NER data characteristics. To address this issue, we propose a Multi-Task Semantic Decomposition Framework via Joint Task-specific Pre-training (MSDP) for few-shot NER. Drawing inspiration from demonstration-based and contrastive learning, we introduce two novel pre-training tasks: Demonstration-based Masked Language Modeling (MLM) and Class Contrastive Discrimination. These tasks effectively incorporate entity boundary information and enhance entity representation in Pre-trained Language Models (PLMs). In the downstream main task, we introduce a multi-task joint optimization framework with the semantic decomposing method, which facilitates the model to integrate two different semantic information for entity classification. Experimental results of two few-shot NER benchmarks demonstrate that MSDP consistently outperforms strong baselines by a large margin. Extensive analyses validate the effectiveness and generalization of MSDP.
Guanting Dong 0001, Zechen Wang, Jinxu Zhao, Daichi Guo, Dayuan Fu, Tingfeng Hui, Keqing He 0001, Xuefeng Li 0002, Liwen Wang 0007, Weiran Xu
CIKM9
2023 Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPT
abstract
Xiaoshuai Song, Keqing He, Pei Wang, Guanting Dong, Yutao Mou, Jingang Wang, Yunsen Xian, Xunliang Cai, Weiran Xu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Xiaoshuai Song, Keqing He 0001, Guanting Dong 0001, Yutao Mou, Jingang Wang, Yunsen Xian, Weiran Xu
EMNLP2
2023 A Prototypical Semantic Decoupling Method via Joint Contrastive Learning for Few-Shot Named Entity Recognition
abstract
Few-shot named entity recognition (NER) aims at identifying named entities based on only few labeled instances. Most existing prototype-based sequence labeling models tend to memorize entity mentions which would be easily confused by close prototypes. In this paper, we proposed a Prototypical Semantic Decoupling method via joint Contrastive learning (PSDC) for few-shot NER. Specifically, we decouple class-specific prototypes and contextual semantic prototypes by two masking strategies to lead the model to focus on two different semantic information for inference. Besides, we further introduce joint contrastive learning objectives to better integrate two kinds of decoupling information and prevent semantic collapse. Experimental results on two few-shot NER benchmarks demonstrate that PSDC consistently outperforms the previous SOTA methods in terms of overall performance. Extensive analysis further validates the effectiveness and generalization of PSDC.
Guanting Dong 0001, Zechen Wang, Liwen Wang 0007, Daichi Guo, Dayuan Fu, Yuxiang Wu, Xuefeng Li 0002, Tingfeng Hui, Keqing He 0001, QiXiang Gao, Weiran Xu
ICASSP10
2023 Revisit Out-Of-Vocabulary Problem For Slot Filling: A Unified Contrastive Framework With Multi-Level Data Augmentations
abstract
In real dialogue scenarios, the existing slot filling model, which tends to memorize entity patterns, has a significantly reduced generalization facing Out-of-Vocabulary (OOV) problems. To address this issue, we propose an OOV robust slot filling model based on multi-level data augmentations to solve the OOV problem from both word and slot perspectives. We present a unified contrastive learning framework, which pull representations of the origin sample and augmentation samples together, to make the model resistant to OOV problems. We evaluate the performance of the model from some specific slots and carefully design test data with OOV word perturbation to further demonstrate the effectiveness of OOV words. Experiments on two datasets show that our approach outperforms the previous sota methods in terms of both OOV slots and words.
Daichi Guo, Guanting Dong 0001, Dayuan Fu, Yuxiang Wu, Tingfeng Hui, Liwen Wang 0007, Xuefeng Li 0002, Zechen Wang, Keqing He 0001, Weiran Xu
ICASSP10
2023 Revisit Input Perturbation Problems for LLMs: A Unified Robustness Evaluation Framework for Noisy Slot Filling Task
Guanting Dong 0001, Jinxu Zhao, Tingfeng Hui, Daichi Guo, Boqi Feng, Yueyan Qiu, Zhuoma Gongque, Keqing He 0001, Zechen Wang, Weiran Xu
NLPCC (1)9
2022 Unified Knowledge Prompt Pre-training for Customer Service Dialogues
abstract
Dialogue bots have been widely applied in customer service scenarios to provide timely and user-friendly experience. These bots must classify the appropriate domain of a dialogue, understand the intent of users, and generate proper responses. Existing dialogue pre-training models are designed only for several dialogue tasks and ignore weakly-supervised expert knowledge in customer service dialogues. In this paper, we propose a novel unified knowledge prompt pre-training framework, UFA (Unified Model F or All Tasks), for customer service dialogues. We formulate all the tasks of customer service dialogues as a unified text-to-text generation task and introduce a knowledge-driven prompt strategy to jointly learn from a mixture of distinct dialogue tasks. We pre-train UFA on a large-scale Chinese customer service corpus collected from practical scenarios and get significant improvements on both natural language understanding (NLU) and natural language generation (NLG) benchmarks.
Keqing He 0001, Jingang Wang, Chaobo Sun, Wei Wu 0014
CIKM1
2022 PSSAT: A Perturbed Semantic Structure Awareness Transferring Method for Perturbation-Robust Slot Filling
abstract
Most existing slot filling models tend to memorize inherent patterns of entities and corresponding contexts from training data. However, these models can lead to system failure or undesirable outputs when being exposed to spoken language perturbation or variation in practice. We propose a perturbed semantic structure awareness transferring method for training perturbation-robust slot filling models. Specifically, we introduce two MLM-based training strategies to respectively learn contextual semantic structure and word distribution from unsupervised language perturbation corpus. Then, we transfer semantic knowledge learned from upstream training procedure into the original samples and filter generated data by consistency processing. These procedures aims to enhance the robustness of slot filling models. Experimental results show that our method consistently outperforms the previous basic methods and gains strong generalization while preventing the model from memorizing inherent patterns of entities and contexts.
Guanting Dong 0001, Daichi Guo, Liwen Wang 0007, Xuefeng Li 0002, Zechen Wang, Keqing He 0001, Jinzheng Zhao, Yi Huang 0017, Junlan Feng, Weiran Xu
COLING7
2022 Generalized Intent Discovery: Learning from Open World Dialogue System
abstract
Traditional intent classification models are based on a pre-defined intent set and only recognize limited in-domain (IND) intent classes. But users may input out-of-domain (OOD) queries in a practical dialogue system. Such OOD queries can provide directions for future improvement. In this paper, we define a new task, Generalized Intent Discovery (GID), which aims to extend an IND intent classifier to an open-world intent set including IND and OOD intents. We hope to simultaneously classify a set of labeled IND intent classes while discovering and recognizing new unlabeled OOD types incrementally. We construct three public datasets for different application scenarios and propose two kinds of frameworks, pipeline-based and end-to-end for future work. Further, we conduct exhaustive experiments and qualitative analysis to comprehend key challenges and provide new guidance for future GID research.
Yutao Mou, Keqing He 0001, Yanan Wu 0002, Jingang Wang, Wei Wu 0014, Yi Huang 0017, Junlan Feng, Weiran Xu
COLING2
2022 Distribution Calibration for Out-of-Domain Detection with Bayesian Approximation
abstract
Out-of-Domain (OOD) detection is a key component in a task-oriented dialog system, which aims to identify whether a query falls outside the predefined supported intent set. Previous softmax-based detection algorithms are proved to be overconfident for OOD samples. In this paper, we analyze overconfident OOD comes from distribution uncertainty due to the mismatch between the training and test distributions, which makes the model can’t confidently make predictions thus probably causes abnormal softmax scores. We propose a Bayesian OOD detection framework to calibrate distribution uncertainty using Monte-Carlo Dropout. Our method is flexible and easily pluggable to existing softmax-based baselines and gains 33.33% OOD F1 improvements with increasing only 0.41% inference time compared to MSP. Further analyses show the effectiveness of Bayesian learning for OOD detection.
Yanan Wu 0002, Zhiyuan Zeng 0002, Keqing He 0001, Yutao Mou, Weiran Xu
COLING3
2022 Watch the Neighbors: A Unified K-Nearest Neighbor Contrastive Learning Framework for OOD Intent Discovery
abstract
Discovering out-of-domain (OOD) intent is important for developing new skills in taskoriented dialogue systems.The key challenges lie in how to transfer prior in-domain (IND) knowledge to OOD clustering, as well as jointly learn OOD representations and cluster assignments.Previous methods suffer from indomain overfitting problem, and there is a natural gap between representation learning and clustering objectives.In this paper, we propose a unified K-nearest neighbor contrastive learning framework to discover OOD intents.Specifically, for IND pre-training stage, we propose a KCL objective to learn inter-class discriminative features, while maintaining intraclass diversity, which alleviates the in-domain overfitting problem.For OOD clustering stage, we propose a KCC method to form compact clusters by mining true hard negative samples, which bridges the gap between clustering and representation learning.Extensive experiments on three benchmark datasets show that our method achieves substantial improvements over the state-of-the-art methods. 1
Yutao Mou, Keqing He 0001, Yanan Wu 0002, Jingang Wang, Wei Wu 0014, Weiran Xu
EMNLP2
2022 UniNL: Aligning Representation Learning with Scoring Function for OOD Detection via Unified Neighborhood Learning
abstract
Detecting out-of-domain (OOD) intents from user queries is essential for avoiding wrong operations in task-oriented dialogue systems.The key challenge is how to distinguish indomain (IND) and OOD intents.Previous methods ignore the alignment between representation learning and scoring function, limiting the OOD detection performance.In this paper, we propose a unified neighborhood learning framework (UniNL) to detect OOD intents.Specifically, we design a K-nearest neighbor contrastive learning (KNCL) objective for representation learning and introduce a KNNbased scoring function for OOD detection.We aim to align representation learning with scoring function.Experiments and analysis on two benchmark datasets show the effectiveness of our method.1
Yutao Mou, Keqing He 0001, Yanan Wu 0002, Jingang Wang, Wei Wu 0014, Weiran Xu
EMNLP3
2022 Revisit Overconfidence for OOD Detection: Reassigned Contrastive Learning with Adaptive Class-dependent Threshold
abstract
Yanan Wu, Keqing He, Yuanmeng Yan, QiXiang Gao, Zhiyuan Zeng, Fujia Zheng, Lulu Zhao, Huixing Jiang, Wei Wu, Weiran Xu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yanan Wu 0002, Keqing He 0001, Yuanmeng Yan, QiXiang Gao, Zhiyuan Zeng 0002, Fujia Zheng, Huixing Jiang, Wei Wu 0014, Weiran Xu
NAACL-HLT2
2022 Domain-Oriented Prefix-Tuning: Towards Efficient and Generalizable Fine-tuning for Zero-Shot Dialogue Summarization
abstract
Lulu Zhao, Fujia Zheng, Weihao Zeng, Keqing He, Weiran Xu, Huixing Jiang, Wei Wu, Yanan Wu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Fujia Zheng, Weihao Zeng 0003, Keqing He 0001, Weiran Xu, Huixing Jiang, Wei Wu 0014, Yanan Wu 0002
NAACL-HLT4
2022 ADPL: Adversarial Prompt-based Domain Adaptation for Dialogue Summarization with Knowledge Disentanglement
abstract
Traditional dialogue summarization models rely on a large-scale manually-labeled corpus, lacking generalization ability to new domains, and domain adaptation from a labeled source domain to an unlabeled target domain is important in practical summarization scenarios. However, existing domain adaptation works in dialogue summarization generally require large-scale pre-training using extensive external data. To explore the lightweight fine-tuning methods, in this paper, we propose an efficient Adversarial Disentangled Prompt Learning (ADPL) model for domain adaptation in dialogue summarization. We introduce three kinds of prompts including domain-invariant prompt (DIP), domain-specific prompt (DSP), and task-oriented prompt (TOP). DIP aims to disentangle and transfer the shared knowledge from the source domain and target domain in an adversarial way, which improves the accuracy of prediction about domain-invariant information and enhances the ability for generalization to new domains. DSP is designed to guide our model to focus on domain-specific knowledge using domain-related features. TOP is to capture task-oriented knowledge to generate high-quality summaries. Instead of fine-tuning the whole pre-trained language model (PLM), we only update the prompt networks but keep PLM fixed. Experimental results on the zero-shot setting show that the novel design of prompts can yield more coherent, faithful, and relevant summaries than baselines using the prefix-tuning, and perform at par with fine-tuning while being more efficient. Overall, our work introduces a prompt-based perspective to the zero-shot learning for dialogue summarization task and provides valuable findings and insights for future research.
Fujia Zheng, Weihao Zeng 0003, Keqing He 0001, Ruotong Geng, Huixing Jiang, Wei Wu 0014, Weiran Xu
SIGIR4
2021 Novel Slot Detection: A Benchmark for Discovering Unknown Slot Types in the Task-Oriented Dialogue System
abstract
Yanan Wu, Zhiyuan Zeng, Keqing He, Hong Xu, Yuanmeng Yan, Huixing Jiang, Weiran Xu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yanan Wu 0002, Zhiyuan Zeng 0002, Keqing He 0001, Hong Xu 0009, Yuanmeng Yan, Huixing Jiang, Weiran Xu
ACL/IJCNLP (1)3
2021 Bridge to Target Domain by Prototypical Contrastive Learning and Label Confusion: Re-explore Zero-Shot Learning for Slot Filling
abstract
Zero-shot cross-domain slot filling alleviates the data dependence in the case of data scarcity in the target domain, which has aroused extensive research.However, as most of the existing methods do not achieve effective knowledge transfer to the target domain, they just fit the distribution of the seen slot and show poor performance on unseen slot in the target domain.To solve this, we propose a novel approach based on prototypical contrastive learning with a dynamic label confusion strategy for zero-shot slot filling.The prototypical contrastive learning aims to reconstruct the semantic constraints of labels, and we introduce the label confusion strategy to establish the label dependence between the source domains and the target domain on-the-fly.Experimental results show that our model achieves significant improvement on the unseen slots, while also set new state-of-the-arts on slot filling task. 1
Liwen Wang 0007, Xuefeng Li 0002, Jiachi Liu, Keqing He 0001, Yuanmeng Yan, Weiran Xu
EMNLP (1)4
2021 Hierarchical Speaker-Aware Sequence-to-Sequence Model for Dialogue Summarization
abstract
Traditional document summarization models cannot handle dialogue summarization tasks perfectly. In situations with multiple speakers and complex personal pronouns referential relationships in the conversation. The predicted summaries of these models are always full of personal pronoun confusion. In this paper, we propose a hierarchical transformer-based model for dialogue summarization. It encodes dialogues from words to utterances and distinguishes the relationships between speakers and their corresponding personal pronouns clearly. In such a from-coarse-to-fine procedure, our model can generate summaries more accurately and relieve the confusion of personal pronouns. Experiments are based on a dialogue summarization dataset SAMsum, and the results show that the proposed model achieved a comparable result against other strong baselines. Empirical experiments have shown that our method can relieve the confusion of personal pronouns in predicted summaries.
Yuejie Lei, Yuanmeng Yan, Zhiyuan Zeng 0002, Keqing He 0001, Weiran Xu
ICASSP4
2021 Adversarial Generative Distance-Based Classifier for Robust Out-of-Domain Detection
abstract
Detecting out-of-domain (OOD) intents is critical in a task-oriented dialog system. Existing methods rely heavily on extensive manually labeled OOD samples and lack robustness. In this paper, we propose an efficient adversarial attack mechanism to augment hard OOD samples and design a novel generative distance-based classifier to detect OOD samples instead of a traditional threshold-based discriminator classifier. Experiments on two public benchmark datasets show that our method can consistently outperform the baselines with a statistically significant margin.
Zhiyuan Zeng 0002, Hong Xu 0009, Keqing He 0001, Yuanmeng Yan, Sihong Liu, Weiran Xu
ICASSP3
2021 Self-training with Masked Supervised Contrastive Loss for Unknown Intents Detection
abstract
The performance of many intent detection approaches will degrade when they meet open-set data because there is out-of-domain (OOD) noise. Some works utilize clean but expensive labeled data to supervise models more robust to the varied environment. However, a large number of labeled samples are scarce and their models no longer change after finishing training. To address this problem, we propose an iterative learning framework that can dynamically improve the model's ability of OOD intent detection and meanwhile continually obtain valuable new data to learn deep discriminative features. Concretely, the model can generate pseudo labels for unlabeled examples by self-training and use the local outlier factor (LOF) algorithm to detect unknown intents. Furthermore, we add mask operation on supervised contrastive loss (SCL) and use the masked-SCL to absorb new data effectively. Experiments on CLINC and SNIPS demonstrate that our proposed method can robustly realize intent detection in the presence of a high proportion of open-set.
Yuanmeng Yan, Keqing He 0001, Sihong Liu, Hong Xu 0009, Weiran Xu
IJCNN3
2021 Dynamically Disentangling Social Bias from Task-Oriented Representations with Adversarial Attack
abstract
Liwen Wang, Yuanmeng Yan, Keqing He, Yanan Wu, Weiran Xu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Liwen Wang 0007, Yuanmeng Yan, Keqing He 0001, Yanan Wu 0002, Weiran Xu
NAACL-HLT3
2021 Adversarial Self-Supervised Learning for Out-of-Domain Detection
abstract
Zhiyuan Zeng, Keqing He, Yuanmeng Yan, Hong Xu, Weiran Xu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Zhiyuan Zeng 0002, Keqing He 0001, Yuanmeng Yan, Hong Xu 0009, Weiran Xu
NAACL-HLT2
2021 From context-aware to knowledge-aware: Boosting OOV tokens recognition in slot tagging with background knowledge
Keqing He 0001, Yuanmeng Yan, Weiran Xu
Neurocomputing1
2020 Learning to Tag OOV Tokens by Integrating Contextual Representation and Background Knowledge
abstract
Neural-based context-aware models for slot tagging have achieved state-of-the-art performance.However, the presence of OOV(outof-vocab) words significantly degrades the performance of neural-based models, especially in a few-shot scenario.In this paper, we propose a novel knowledge-enhanced slot tagging model to integrate contextual representation of input text and the large-scale lexical background knowledge.Besides, we use multilevel graph attention to explicitly model lexical relations.The experiments show that our proposed knowledge integration mechanism achieves consistent improvements across settings with different sizes of training data on two public benchmark datasets.
Keqing He 0001, Yuanmeng Yan, Weiran Xu
ACL1
2020 Syntactic Graph Convolutional Network for Spoken Language Understanding
abstract
Slot filling and intent detection are two major tasks for spoken language understanding.In most existing work, these two tasks are built as joint models with multi-task learning with no consideration of prior linguistic knowledge.In this paper, we propose a novel joint model that applies a graph convolutional network over dependency trees to integrate the syntactic structure for learning slot filling and intent detection jointly.Experimental results show that our proposed model achieves state-of-the-art performance on two public benchmark datasets and outperforms existing work.At last, we apply the BERT model to further improve the performance on both slot filling and intent detection.
Keqing He 0001, Shuyu Lei, Yushu Yang, Huixing Jiang
COLING1
2020 Contrastive Zero-Shot Learning for Cross-Domain Slot Filling with Adversarial Attack
abstract
Zero-shot slot filling has widely arisen to cope with data scarcity in target domains.However, previous approaches often ignore constraints between slot value representation and related slot description representation in the latent space and lack enough model robustness.In this paper, we propose a Contrastive Zero-Shot Learning with Adversarial Attack (CZSL-Adv) method for the cross-domain slot filling.The contrastive loss aims to map slot value contextual representations to the corresponding slot description representations.And we introduce an adversarial attack training strategy to improve model robustness.Experimental results show that our model significantly outperforms state-of-the-art baselines under both zero-shot and few-shot settings.
Keqing He 0001, Jinchao Zhang 0001, Yuanmeng Yan, Weiran Xu, Cheng Niu, Jie Zhou 0016
COLING1
2020 A Deep Generative Distance-Based Classifier for Out-of-Domain Detection with Mahalanobis Space
abstract
Detecting out-of-domain (OOD) input intents is critical in the task-oriented dialog system. Different from most existing methods that rely heavily on manually labeled OOD samples, we focus on the unsupervised OOD detection scenario where there are no labeled OOD samples except for labeled in-domain data. In this paper, we propose a simple but strong generative distance-based classifier to detect OOD samples. We estimate the class-conditional distribution on feature spaces of DNNs via Gaussian discriminant analysis (GDA) to avoid over-confidence problems. And we use two distance functions, Euclidean and Mahalanobis distances, to measure the confidence score of whether a test sample belongs to OOD. Experiments on four benchmark datasets show that our method can consistently outperform the baselines.
Hong Xu 0009, Keqing He 0001, Yuanmeng Yan, Sihong Liu, Weiran Xu
COLING2
2020 Adversarial Semantic Decoupling for Recognizing Open-Vocabulary Slots
abstract
Open-vocabulary slots, such as file name, album name, or schedule title, significantly degrade the performance of neural-based slot filling models since these slots can take on values from a virtually unlimited set and have no semantic restriction nor a length limit.In this paper, we propose a robust adversarial model-agnostic slot filling method that explicitly decouples local semantics inherent in open-vocabulary slot words from the global context.We aim to depart entangled contextual semantics and focus more on the holistic context at the level of the whole sentence.Experiments on two public datasets show that our method consistently outperforms other methods with a statistically significant margin on all the open-vocabulary slots without deteriorating the performance of normal slots. *The first two authors contribute equally.Weiran Xu is the corresponding author.
Yuanmeng Yan, Keqing He 0001, Hong Xu 0009, Sihong Liu, Weiran Xu
EMNLP (1)2
2020 Adversarial Cross-Lingual Transfer Learning for Slot Tagging of Low-Resource Languages
abstract
Slot tagging is a key component in a task-oriented dialogue system. Conversational agents need to understand human input by training on large amounts of annotated data. However, most human languages are low-resource and lack annotated training data for slot tagging task. Therefore, we aim to leverage cross-lingual transfer learning from high-resource languages to low-resource ones. In this paper, we propose an adversarial cross-lingual transfer model with multi-level language shared and specific knowledge to improve the slot tagging task of low-resource languages. Our method explicitly separates the model into the language-shared part and language-specific part to transfer language-independent knowledge. To refine shared knowledge in the latent space, we add a language discriminator and employ adversarial training to reinforce feature separation. Besides, we adopt a novel multi-level feature transfer in an incremental and progressive way to acquire multi-granularity shared knowledge. To mitigate the discrepancies between the feature distributions of language specific and shared knowledge, we propose the neural adapters to fuse features from different sources. Experiments show that our proposed model consistently outperforms monolingual baseline with a statistically significant margin up to 2.09%, even higher improvement of 12.21% in the zero-shot setting. Further analysis demonstrates that our method could effectively alleviate data scarcity of low-resource languages.
Keqing He 0001, Yuanmeng Yan, Weiran Xu
IJCNN1
2020 Learning Label-Relational Output Structure for Adaptive Sequence Labeling
abstract
Sequence labeling is a fundamental task of natural language understanding. Recent neural models for sequence labeling task achieve significant success with the availability of sufficient training data. However, in practical scenarios, entity types to be annotated even in the same domain are continuously evolving. To transfer knowledge from the source model pre-trained on previously annotated data, we propose an approach which learns label-relational output structure to explicitly capturing label correlations in the latent space. Additionally, we construct the target-to-source interaction between the source model MSand the target model MTand apply a gate mechanism to control how much information in MSand MTshould be passed down. Experiments show that our method consistently outperforms the state-of-the-art methods with a statistically significant margin and effectively facilitates to recognize rare new entities in the target data especially.
Keqing He 0001, Yuanmeng Yan, Hong Xu 0009, Sihong Liu, Weiran Xu
IJCNN1