Yuanmeng Yan

dblp:259/0285 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
13since 2021 · last 2022
0000-0002-0400-4522ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2022 Revisit Overconfidence for OOD Detection: Reassigned Contrastive Learning with Adaptive Class-dependent Threshold
abstract
Yanan Wu, Keqing He, Yuanmeng Yan, QiXiang Gao, Zhiyuan Zeng, Fujia Zheng, Lulu Zhao, Huixing Jiang, Wei Wu, Weiran Xu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yanan Wu 0002, Keqing He 0001, Yuanmeng Yan, QiXiang Gao, Zhiyuan Zeng 0002, Fujia Zheng, Huixing Jiang, Wei Wu 0014, Weiran Xu
NAACL-HLT3
2022 S2QL: Retrieval Augmented Zero-Shot Question Answering over Knowledge Graph
Daoguang Zan, Yuanmeng Yan, Wei Wu 0014, Bei Guan, Yongji Wang 0002
PAKDD (3)4
2021 Novel Slot Detection: A Benchmark for Discovering Unknown Slot Types in the Task-Oriented Dialogue System
abstract
Yanan Wu, Zhiyuan Zeng, Keqing He, Hong Xu, Yuanmeng Yan, Huixing Jiang, Weiran Xu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yanan Wu 0002, Zhiyuan Zeng 0002, Keqing He 0001, Hong Xu 0009, Yuanmeng Yan, Huixing Jiang, Weiran Xu
ACL/IJCNLP (1)5
2021 ConSERT: A Contrastive Framework for Self-Supervised Sentence Representation Transfer
abstract
Yuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang, Wei Wu, Weiran Xu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yuanmeng Yan, Wei Wu 0014, Weiran Xu
ACL/IJCNLP (1)1
2021 Bridge to Target Domain by Prototypical Contrastive Learning and Label Confusion: Re-explore Zero-Shot Learning for Slot Filling
abstract
Zero-shot cross-domain slot filling alleviates the data dependence in the case of data scarcity in the target domain, which has aroused extensive research.However, as most of the existing methods do not achieve effective knowledge transfer to the target domain, they just fit the distribution of the seen slot and show poor performance on unseen slot in the target domain.To solve this, we propose a novel approach based on prototypical contrastive learning with a dynamic label confusion strategy for zero-shot slot filling.The prototypical contrastive learning aims to reconstruct the semantic constraints of labels, and we introduce the label confusion strategy to establish the label dependence between the source domains and the target domain on-the-fly.Experimental results show that our model achieves significant improvement on the unseen slots, while also set new state-of-the-arts on slot filling task. 1
Liwen Wang 0007, Xuefeng Li 0002, Jiachi Liu, Keqing He 0001, Yuanmeng Yan, Weiran Xu
EMNLP (1)5
2021 Large-Scale Relation Learning for Question Answering over Knowledge Bases with Pre-trained Language Models
abstract
The key challenge of question answering over knowledge bases (KBQA) is the inconsistency between the natural language questions and the reasoning paths in the knowledge base (KB).Recent graph-based KBQA methods are good at grasping the topological structure of the graph but often ignore the textual information carried by the nodes and edges.Meanwhile, pre-trained language models learn massive open-world knowledge from the large corpus, but it is in the natural language form and not structured.To bridge the gap between the natural language and the structured KB, we propose three relation learning tasks for BERTbased KBQA, including relation extraction, relation matching, and relation reasoning.By relation-augmented training, the model learns to align the natural language expressions to the relations in the KB as well as reason over the missing connections in the KB.Experiments on WebQSP show that our method consistently outperforms other baselines, especially when the KB is incomplete.
Yuanmeng Yan, Daoguang Zan, Wei Wu 0014, Weiran Xu
EMNLP (1)1
2021 Hierarchical Speaker-Aware Sequence-to-Sequence Model for Dialogue Summarization
abstract
Traditional document summarization models cannot handle dialogue summarization tasks perfectly. In situations with multiple speakers and complex personal pronouns referential relationships in the conversation. The predicted summaries of these models are always full of personal pronoun confusion. In this paper, we propose a hierarchical transformer-based model for dialogue summarization. It encodes dialogues from words to utterances and distinguishes the relationships between speakers and their corresponding personal pronouns clearly. In such a from-coarse-to-fine procedure, our model can generate summaries more accurately and relieve the confusion of personal pronouns. Experiments are based on a dialogue summarization dataset SAMsum, and the results show that the proposed model achieved a comparable result against other strong baselines. Empirical experiments have shown that our method can relieve the confusion of personal pronouns in predicted summaries.
Yuejie Lei, Yuanmeng Yan, Zhiyuan Zeng 0002, Keqing He 0001, Weiran Xu
ICASSP2
2021 Boosting Low-Resource Intent Detection with in-Scope Prototypical Networks
abstract
Identifying intentions from users can help improve the response quality of task-oriented dialogue systems. How to use only limited labeled in-domain (ID) examples for zero-shot unknown intent detection and few-shot ID classification is a more challenging task in spoken language understanding. Existing related methods heavily rely upon the multi-domain datasets containing large-scale independent source domains for meta-training. In this paper, we propose a universal In-scope Prototypical Networks for low-resource intent detection to be general to dialogue meta-train datasets lacking widely-varying domains, which focuses on the scope of episodic intent classes to construct meta-task dynamically. Also, we introduce loss with margin principle to better distinguish samples. Experiments on two benchmark datasets show that our model consistently outperforms other baselines on zero-shot unknown intent detection without deteriorating the competitive performance on few-shot ID classification.
Hongzhan Lin 0001, Yuanmeng Yan, Guang Chen 0003
ICASSP2
2021 Adversarial Generative Distance-Based Classifier for Robust Out-of-Domain Detection
abstract
Detecting out-of-domain (OOD) intents is critical in a task-oriented dialog system. Existing methods rely heavily on extensive manually labeled OOD samples and lack robustness. In this paper, we propose an efficient adversarial attack mechanism to augment hard OOD samples and design a novel generative distance-based classifier to detect OOD samples instead of a traditional threshold-based discriminator classifier. Experiments on two public benchmark datasets show that our method can consistently outperform the baselines with a statistically significant margin.
Zhiyuan Zeng 0002, Hong Xu 0009, Keqing He 0001, Yuanmeng Yan, Sihong Liu, Weiran Xu
ICASSP4
2021 Self-training with Masked Supervised Contrastive Loss for Unknown Intents Detection
abstract
The performance of many intent detection approaches will degrade when they meet open-set data because there is out-of-domain (OOD) noise. Some works utilize clean but expensive labeled data to supervise models more robust to the varied environment. However, a large number of labeled samples are scarce and their models no longer change after finishing training. To address this problem, we propose an iterative learning framework that can dynamically improve the model's ability of OOD intent detection and meanwhile continually obtain valuable new data to learn deep discriminative features. Concretely, the model can generate pseudo labels for unlabeled examples by self-training and use the local outlier factor (LOF) algorithm to detect unknown intents. Furthermore, we add mask operation on supervised contrastive loss (SCL) and use the masked-SCL to absorb new data effectively. Experiments on CLINC and SNIPS demonstrate that our proposed method can robustly realize intent detection in the presence of a high proportion of open-set.
Yuanmeng Yan, Keqing He 0001, Sihong Liu, Hong Xu 0009, Weiran Xu
IJCNN2
2021 Dynamically Disentangling Social Bias from Task-Oriented Representations with Adversarial Attack
abstract
Liwen Wang, Yuanmeng Yan, Keqing He, Yanan Wu, Weiran Xu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Liwen Wang 0007, Yuanmeng Yan, Keqing He 0001, Yanan Wu 0002, Weiran Xu
NAACL-HLT2
2021 Adversarial Self-Supervised Learning for Out-of-Domain Detection
abstract
Zhiyuan Zeng, Keqing He, Yuanmeng Yan, Hong Xu, Weiran Xu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Zhiyuan Zeng 0002, Keqing He 0001, Yuanmeng Yan, Hong Xu 0009, Weiran Xu
NAACL-HLT3
2021 From context-aware to knowledge-aware: Boosting OOV tokens recognition in slot tagging with background knowledge
Keqing He 0001, Yuanmeng Yan, Weiran Xu
Neurocomputing2
2020 Learning to Tag OOV Tokens by Integrating Contextual Representation and Background Knowledge
abstract
Neural-based context-aware models for slot tagging have achieved state-of-the-art performance.However, the presence of OOV(outof-vocab) words significantly degrades the performance of neural-based models, especially in a few-shot scenario.In this paper, we propose a novel knowledge-enhanced slot tagging model to integrate contextual representation of input text and the large-scale lexical background knowledge.Besides, we use multilevel graph attention to explicitly model lexical relations.The experiments show that our proposed knowledge integration mechanism achieves consistent improvements across settings with different sizes of training data on two public benchmark datasets.
Keqing He 0001, Yuanmeng Yan, Weiran Xu
ACL2
2020 Contrastive Zero-Shot Learning for Cross-Domain Slot Filling with Adversarial Attack
abstract
Zero-shot slot filling has widely arisen to cope with data scarcity in target domains.However, previous approaches often ignore constraints between slot value representation and related slot description representation in the latent space and lack enough model robustness.In this paper, we propose a Contrastive Zero-Shot Learning with Adversarial Attack (CZSL-Adv) method for the cross-domain slot filling.The contrastive loss aims to map slot value contextual representations to the corresponding slot description representations.And we introduce an adversarial attack training strategy to improve model robustness.Experimental results show that our model significantly outperforms state-of-the-art baselines under both zero-shot and few-shot settings.
Keqing He 0001, Jinchao Zhang 0001, Yuanmeng Yan, Weiran Xu, Cheng Niu, Jie Zhou 0016
COLING3
2020 A Deep Generative Distance-Based Classifier for Out-of-Domain Detection with Mahalanobis Space
abstract
Detecting out-of-domain (OOD) input intents is critical in the task-oriented dialog system. Different from most existing methods that rely heavily on manually labeled OOD samples, we focus on the unsupervised OOD detection scenario where there are no labeled OOD samples except for labeled in-domain data. In this paper, we propose a simple but strong generative distance-based classifier to detect OOD samples. We estimate the class-conditional distribution on feature spaces of DNNs via Gaussian discriminant analysis (GDA) to avoid over-confidence problems. And we use two distance functions, Euclidean and Mahalanobis distances, to measure the confidence score of whether a test sample belongs to OOD. Experiments on four benchmark datasets show that our method can consistently outperform the baselines.
Hong Xu 0009, Keqing He 0001, Yuanmeng Yan, Sihong Liu, Weiran Xu
COLING3
2020 Adversarial Semantic Decoupling for Recognizing Open-Vocabulary Slots
abstract
Open-vocabulary slots, such as file name, album name, or schedule title, significantly degrade the performance of neural-based slot filling models since these slots can take on values from a virtually unlimited set and have no semantic restriction nor a length limit.In this paper, we propose a robust adversarial model-agnostic slot filling method that explicitly decouples local semantics inherent in open-vocabulary slot words from the global context.We aim to depart entangled contextual semantics and focus more on the holistic context at the level of the whole sentence.Experiments on two public datasets show that our method consistently outperforms other methods with a statistically significant margin on all the open-vocabulary slots without deteriorating the performance of normal slots. *The first two authors contribute equally.Weiran Xu is the corresponding author.
Yuanmeng Yan, Keqing He 0001, Hong Xu 0009, Sihong Liu, Weiran Xu
EMNLP (1)1
2020 Adversarial Cross-Lingual Transfer Learning for Slot Tagging of Low-Resource Languages
abstract
Slot tagging is a key component in a task-oriented dialogue system. Conversational agents need to understand human input by training on large amounts of annotated data. However, most human languages are low-resource and lack annotated training data for slot tagging task. Therefore, we aim to leverage cross-lingual transfer learning from high-resource languages to low-resource ones. In this paper, we propose an adversarial cross-lingual transfer model with multi-level language shared and specific knowledge to improve the slot tagging task of low-resource languages. Our method explicitly separates the model into the language-shared part and language-specific part to transfer language-independent knowledge. To refine shared knowledge in the latent space, we add a language discriminator and employ adversarial training to reinforce feature separation. Besides, we adopt a novel multi-level feature transfer in an incremental and progressive way to acquire multi-granularity shared knowledge. To mitigate the discrepancies between the feature distributions of language specific and shared knowledge, we propose the neural adapters to fuse features from different sources. Experiments show that our proposed model consistently outperforms monolingual baseline with a statistically significant margin up to 2.09%, even higher improvement of 12.21% in the zero-shot setting. Further analysis demonstrates that our method could effectively alleviate data scarcity of low-resource languages.
Keqing He 0001, Yuanmeng Yan, Weiran Xu
IJCNN2
2020 Learning Label-Relational Output Structure for Adaptive Sequence Labeling
abstract
Sequence labeling is a fundamental task of natural language understanding. Recent neural models for sequence labeling task achieve significant success with the availability of sufficient training data. However, in practical scenarios, entity types to be annotated even in the same domain are continuously evolving. To transfer knowledge from the source model pre-trained on previously annotated data, we propose an approach which learns label-relational output structure to explicitly capturing label correlations in the latent space. Additionally, we construct the target-to-source interaction between the source model MSand the target model MTand apply a gate mechanism to control how much information in MSand MTshould be passed down. Experiments show that our method consistently outperforms the state-of-the-art methods with a statistically significant margin and effectively facilitates to recognize rare new entities in the target data especially.
Keqing He 0001, Yuanmeng Yan, Hong Xu 0009, Sihong Liu, Weiran Xu
IJCNN2