Yinhe Zheng

dblp:228/5534 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
15since 2021 · last 2024
0000-0002-7029-5671ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2024 Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models
abstract
Longze Chen, Ziqiang Liu, Wanwei He, Yinhe Zheng, Hao Sun, Yunshui Li, Run Luo, Min Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Longze Chen, Wanwei He, Yinhe Zheng, Yunshui Li, Run Luo, Min Yang 0007
ACL (1)4
2024 Out-of-Domain Intent Detection Considering Multi-Turn Dialogue Contexts
abstract
Out-of-Domain (OOD) intent detection is vital for practical dialogue systems, and it usually requires considering multi-turn dialogue contexts. However, most previous OOD intent detection approaches are limited to single dialogue turns. In this paper, we introduce a context-aware OOD intent detection (Caro) framework to model multi-turn contexts in OOD intent detection tasks. Specifically, we follow the information bottleneck principle to extract robust representations from multi-turn dialogue contexts. Two different views are constructed for each input sample and the superfluous information not related to intent detection is removed using a multi-view information bottleneck loss. Moreover, we also explore utilizing unlabeled data in Caro. A two-stage training process is introduced to mine OOD samples from these unlabeled data, and these OOD samples are used to train the resulting model with a bootstrapping approach. Comprehensive experiments demonstrate that Caro establishes state-of-the-art performances on multi-turn OOD detection tasks by improving the F1-OOD score of over 29% compared to the previous best method.
Hao Lang, Yinhe Zheng, Binyuan Hui, Fei Huang 0002, Yongbin Li 0001
LREC/COLING2
2023 Long-Tailed Question Answering in an Open World
abstract
Real-world data often have an open long-tailed distribution, and building a unified QA model supporting various tasks is vital for practical QA applications.However, it is non-trivial to extend previous QA approaches since they either require access to seen tasks of adequate samples or do not explicitly model samples from unseen tasks.In this paper, we define Open Long-Tailed QA (OLTQA) as learning from long-tailed distributed data and optimizing performance over seen and unseen QA tasks.We propose an OLTQA model that encourages knowledge sharing between head, tail and unseen tasks, and explicitly mines knowledge from a large pre-trained language model (LM).Specifically, we organize our model through a pool of fine-grained components and dynamically combine these components for an input to facilitate knowledge sharing.A retrieve-then-rerank frame is further introduced to select in-context examples, which guild the LM to generate text that express knowledge for QA tasks.Moreover, a twostage training approach is introduced to pretrain the framework by knowledge distillation (KD) from the LM and then jointly train the frame and a QA model through an adaptive mutual KD method.On a large-scale OLTQA dataset we curate from 43 existing QA datasets, our model consistently outperforms the stateof-the-art.We release the code and data
Hao Lang, Yinhe Zheng, Fei Huang 0002, Yongbin Li 0001
ACL (1)3
2023 Empathetic Response Generation via Emotion Cause Transition Graph
abstract
Empathetic dialogue is a human-like behavior that requires the perception of both affective factors (e.g., emotion status) and cognitive factors (e.g., cause of the emotion). Besides concerning emotion status in early work, the latest approaches study emotion causes in empathetic dialogue. These approaches focus on understanding and duplicating emotion causes in the context to show empathy for the speaker. However, instead of only repeating the contextual causes, the real empathic response often demonstrate a logical and emotion-centered transition from the causes in the context to those in the responses. In this work, we propose an emotion cause transition graph to explicitly model the natural transition of emotion causes between two adjacent turns in empathetic dialogue. With this graph, the concept words of the emotion causes in the next turn can be predicted and used by a specifically designed concept-aware decoder to generate the empathic response. Automatic and human experimental results on the benchmark dataset demonstrate that our method produces more empathetic, coherent, informative, and specific responses than existing models.
Yushan Qian, Ting-En Lin, Yinhe Zheng, Yuexian Hou, Yuchuan Wu
ICASSP4
2022 GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy Injection
abstract
Pre-trained models have proved to be powerful in enhancing task-oriented dialog systems. However, current pre-training methods mainly focus on enhancing dialog understanding and generation tasks while neglecting the exploitation of dialog policy. In this paper, we propose GALAXY, a novel pre-trained dialog model that explicitly learns dialog policy from limited labeled dialogs and large-scale unlabeled dialog corpora via semi-supervised learning. Specifically, we introduce a dialog act prediction task for policy optimization during pre-training and employ a consistency regularization term to refine the learned representation with the help of unlabeled dialogs. We also implement a gating mechanism to weigh suitable unlabeled dialog samples. Empirical results show that GALAXY substantially improves the performance of task-oriented dialog systems, and achieves new state-of-the-art results on benchmark datasets: In-Car, MultiWOZ2.0 and MultiWOZ2.1, improving their end-to-end combined scores by 2.5, 5.3 and 5.5 points, respectively. We also show that GALAXY has a stronger few-shot ability than existing models under various low-resource settings. For reproducibility, we release the code and data at https://github.com/siat-nlp/GALAXY.
Wanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu, Zheng Cao 0003, Dermot Liu, Min Yang 0007, Fei Huang 0002, Luo Si, Jian Sun 0021, Yongbin Li 0001
AAAI3
2022 Improving Meta-learning for Low-resource Text Classification and Generation via Memory Imitation
abstract
Yingxiu Zhao, Zhiliang Tian, Huaxiu Yao, Yinhe Zheng, Dongkyu Lee, Yiping Song, Jian Sun, Nevin Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yingxiu Zhao, Zhiliang Tian, Huaxiu Yao, Yinhe Zheng, Yiping Song, Jian Sun 0021, Nevin Lianwen Zhang
ACL (1)4
2022 LayerConnect: Hypernetwork-Assisted Inter-Layer Connector to Enhance Parameter Efficiency
abstract
Pre-trained Language Models (PLMs) are the cornerstone of the modern Natural Language Processing (NLP). However, as PLMs become heavier, fine tuning all their parameters loses their efficiency. Existing parameter-efficient methods generally focus on reducing the trainable parameters in PLMs but neglect the inference speed, which limits the ability to deploy PLMs. In this paper, we propose LayerConnect (hypernetwork-assisted inter-layer connectors) to enhance inference efficiency. Specifically, a light-weight connector with a linear structure is inserted between two Transformer layers, and the parameters inside each connector are tuned by a hypernetwork comprising an interpolator and a down-sampler. We perform extensive experiments on the widely used the GLUE benchmark. The experimental results verify the inference efficiency of our model. Compared to Adapter, our model parameters are reduced to approximately 11.75%, while the performance degradation is kept to less than 5% (2.5 points on average).
Haoxiang Shi, Jiaan Wang, Cen Wang, Yinhe Zheng, Tetsuya Sakai
COLING5
2022 Estimating Soft Labels for Out-of-Domain Intent Detection
abstract
Out-of-Domain (OOD) intent detection is important for practical dialog systems.To alleviate the issue of lacking OOD training samples, some works propose synthesizing pseudo OOD samples and directly assigning one-hot OOD labels to these pseudo samples.However, these one-hot labels introduce noises to the training process because some "hard" pseudo OOD samples may coincide with In-Domain (IND) intents.In this paper, we propose an adaptive soft pseudo labeling (ASoul) method that can estimate soft labels for pseudo OOD samples when training OOD detectors.Semantic connections between pseudo OOD samples and IND intents are captured using an embedding graph.A co-training framework is further introduced to produce resulting soft labels following the smoothness assumption, i.e., close samples are likely to have similar labels.Extensive experiments on three benchmark datasets show that ASoul consistently improves the OOD detection performance and outperforms various competitive baselines.
Hao Lang, Yinhe Zheng, Jian Sun 0021, Fei Huang 0002, Luo Si, Yongbin Li 0001
EMNLP2
2022 Prompt Conditioned VAE: Enhancing Generative Replay for Lifelong Learning in Task-Oriented Dialogue
abstract
Lifelong learning (LL) is vital for advanced task-oriented dialogue (ToD) systems.To address the catastrophic forgetting issue of LL, generative replay methods are widely employed to consolidate past knowledge with generated pseudo samples.However, most existing generative replay methods use only a single taskspecific token to control their models.This scheme is usually not strong enough to constrain the generative model due to insufficient information involved.In this paper, we propose a novel method, prompt conditioned VAE for lifelong learning (PCLL), to enhance generative replay by incorporating tasks' statistics.PCLL captures task-specific distributions with a conditional variational autoencoder, conditioned on natural language prompts to guide the pseudo-sample generation.Moreover, it leverages a distillation process to further consolidate past knowledge by alleviating the noise in pseudo samples.Experiments on natural language understanding tasks of ToD systems demonstrate that PCLL significantly outperforms competitive baselines in building lifelong learning models.We release the code and data at GitHub.
Yingxiu Zhao, Yinhe Zheng, Zhiliang Tian, Jian Sun 0021, Nevin Lianwen Zhang
EMNLP2
2022 CDConv: A Benchmark for Contradiction Detection in Chinese Conversations
abstract
Chujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng, Zhen Guo, Wenquan Wu, Zheng-Yu Niu, Hua Wu, Minlie Huang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Chujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng, Wenquan Wu, Zhengyu Niu, Hua Wu 0003, Minlie Huang
EMNLP3
2022 MMChat: Multi-Modal Chat Dataset on Social Media
abstract
Incorporating multi-modal contexts in conversation is an important step for developing more engaging dialogue systems. In this work, we explore this direction by introducing MMChat: a large scale Chinese multi-modal dialogue corpus (32.4M raw dialogues and 120.84K filtered dialogues). Unlike previous corpora that are crowd-sourced or collected from fictitious movies, MMChat contains image-grounded dialogues collected from real conversations on social media, in which the sparsity issue is observed. Specifically, image-initiated dialogues in common communications may deviate to some non-image-grounded topics as the conversation proceeds. To better investigate this issue, we manually annotate 100K dialogues from MMChat and further filter the corpus accordingly, which yields MMChat-hf. We develop a benchmark model to address the sparsity issue in dialogue generation tasks by adapting the attention routing mechanism on image features. Experiments demonstrate the usefulness of incorporating image features and the effectiveness in handling the sparsity of image features.
Yinhe Zheng, Guanyi Chen, Jian Sun 0021
LREC1
2021 Stylized Dialogue Response Generation Using Stylized Unpaired Texts
abstract
Generating stylized responses is essential to build intelligent and engaging dialogue systems. However, this task is far from well-explored due to the difficulties of rendering a particular style in coherent responses, especially when the target style is embedded only in unpaired texts that cannot be directly used to train the dialogue model. This paper proposes a stylized dialogue generation method that can capture stylistic features embedded in unpaired texts. Specifically, our method can produce dialogue responses that are both coherent to the given context and conform to the target style. In this study, an inverse dialogue model is first introduced to predict possible posts for the input responses. Then this inverse model is used to generate stylized pseudo dialogue pairs based on these stylized unpaired texts. Further, these pseudo pairs are employed to train the stylized dialogue model with a joint training process. A style routing approach is proposed to intensify stylistic features in the decoder. Automatic and manual evaluations on two datasets demonstrate that our method outperforms competitive baselines in producing coherent and style-intensive dialogue responses.
Yinhe Zheng, Zikai Chen, Shilei Huang, Xiaoxi Mao, Minlie Huang
AAAI1
2021 Diversifying Dialog Generation via Adaptive Label Smoothing
abstract
Yida Wang, Yinhe Zheng, Yong Jiang, Minlie Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yida Wang 0009, Yinhe Zheng, Yong Jiang 0001, Minlie Huang
ACL/IJCNLP (1)2
2021 Transferable Persona-Grounded Dialogues via Grounded Minimal Edits
abstract
Grounded dialogue models generate responses that are grounded on certain concepts.Limited by the distribution of grounded dialogue data, models trained on such data face the transferability challenges in terms of the data distribution and the type of grounded concepts.To address the challenges, we propose the grounded minimal editing framework, which minimally edits existing responses to be grounded on the given concept.Focusing on personas, we propose Grounded Minimal Editor (GME), which learns to edit by disentangling and recombining persona-related and persona-agnostic parts of the response.To evaluate persona-grounded minimal editing, we present the PERSONAMI-NEDIT dataset, and experimental results show that GME outperforms competitive baselines by a large margin.To evaluate the transferability, we experiment on the test set of BLEND-EDSKILLTALK and show that GME can edit dialogue models' responses to largely improve their persona consistency while preserving the use of knowledge and empathy. 1
Chen Henry Wu, Yinhe Zheng, Xiaoxi Mao, Minlie Huang
EMNLP (1)2
2021 DS-UI: Dual-Supervised Mixture of Gaussian Mixture Models for Uncertainty Inference in Image Recognition
abstract
This paper proposes a dual-supervised uncertainty inference (DS-UI) framework for improving Bayesian estimation-based UI in DNN-based image recognition. In the DS-UI, we combine the classifier of a DNN, i.e., the last fully-connected (FC) layer, with a mixture of Gaussian mixture models (MoGMM) to obtain an MoGMM-FC layer. Unlike existing UI methods for DNNs, which only calculate the means or modes of the DNN outputs' distributions, the proposed MoGMM-FC layer acts as a probabilistic interpreter for the features that are inputs of the classifier to directly calculate the probabilities of them for the DS-UI. In addition, we propose a dual-supervised stochastic gradient-based variational Bayes (DS-SGVB) algorithm for the MoGMM-FC layer optimization. Unlike conventional SGVB and optimization algorithms in other UI methods, the DS-SGVB not only models the samples in the specific class for each Gaussian mixture model (GMM) in the MoGMM, but also considers the negative samples from other classes for the GMM to reduce the intra-class distances and enlarge the inter-class margins simultaneously for enhancing the learning ability of the MoGMM-FC layer in the DS-UI. Experimental results show the DS-UI outperforms the state-of-the-art UI methods in misclassification detection. We further evaluate the DS-UI in open-set out-of-domain/-distribution detection and find statistically significant improvements. Visualizations of the feature spaces demonstrate the superiority of the DS-UI. Codes are available at https://github.com/PRIS-CV/DS-UI.
Jiyang Xie 0001, Zhanyu Ma, Jing-Hao Xue, Guoqiang Zhang 0003, Yinhe Zheng, Jun Guo 0002
IEEE Trans. Image Process.6
2020 A Pre-Training Based Personalized Dialogue Generation Model with Persona-Sparse Data
abstract
Endowing dialogue systems with personas is essential to deliver more human-like conversations. However, this problem is still far from well explored due to the difficulties of both embodying personalities in natural languages and the persona sparsity issue observed in most dialogue corpora. This paper proposes a pre-training based personalized dialogue model that can generate coherent responses using persona-sparse dialogue data. In this method, a pre-trained language model is used to initialize an encoder and decoder, and personal attribute embeddings are devised to model richer dialogue contexts by encoding speakers' personas together with dialogue histories. Further, to incorporate the target persona in the decoding process and to balance its contribution, an attention routing structure is devised in the decoder to merge features extracted from the target persona and dialogue contexts using dynamically predicted weights. Our model can utilize persona-sparse dialogues in a unified manner during the training process, and can also control the amount of persona-related features to exhibit during the inference process. Both automatic and manual evaluation demonstrates that the proposed model outperforms state-of-the-art methods for generating more coherent and persona consistent responses with persona-sparse data.
Yinhe Zheng, Minlie Huang, Xiaoxi Mao
AAAI1
2020 Dialogue Distillation: Open-Domain Dialogue Augmentation Using Unpaired Data
abstract
Recent advances in open-domain dialogue systems rely on the success of neural models that are trained on large-scale data.However, collecting large-scale dialogue data is usually time-consuming and labor-intensive.To address this data dilemma, we propose a novel data augmentation method for training opendomain dialogue models by utilizing unpaired data.Specifically, a data-level distillation process is first proposed to construct augmented dialogues where both post and response are retrieved from the unpaired data.A ranking module is employed to filter out low-quality dialogues.Further, a model-level distillation process is employed to distill a teacher model trained on high-quality paired data to augmented dialogue pairs, thereby preventing dialogue models from being affected by the noise in the augmented data.Automatic and manual evaluation indicates that our method can produce high-quality dialogue pairs with diverse contents, and the proposed data-level and model-level dialogue distillation can improve the performance of competitive baselines.
Yinhe Zheng, Jianzhi Shao, Xiaoxi Mao, Yadong Xi, Minlie Huang
EMNLP (1)2
2020 Listener's Social Identity Matters in Personalised Response Generation
abstract
Personalised response generation enables generating human-like responses by means of assigning the generator a social identity.However, pragmatics theory suggests that human beings adjust the way of speaking based on not only who they are but also whom they are talking to.In other words, when modelling personalised dialogues, it might be favourable if we also take the listener's social identity into consideration.To validate this idea, we use gender as a typical example of a social variable to investigate how the listener's identity influences the language used in Chinese dialogues on social media.Also, we build personalised generators.The experiment results demonstrate that the listener's identity indeed matters in the language use of responses and that the response generator can capture such differences in language use.More interestingly, by additionally modelling the listener's identity, the personalised response generator performs better in its own identity.
Guanyi Chen, Yinhe Zheng, Yupei Du
INLG2
2020 A Large-Scale Chinese Short-Text Conversation Dataset
Yida Wang 0009, Pei Ke, Yinhe Zheng, Kaili Huang, Yong Jiang 0001, Xiaoyan Zhu 0001, Minlie Huang
NLPCC (1)3
2020 Out-of-Domain Detection for Natural Language Understanding in Dialog Systems
abstract
Natural Language Understanding (NLU) is a vital component of dialogue systems, and its ability to detect Out-of-Domain (OOD) inputs is critical in practical applications, since the acceptance of the OOD input that is unsupported by the current system may lead to catastrophic failure. However, most existing OOD detection methods rely heavily on manually labeled OOD samples and cannot take full advantage of unlabeled data. This limits the feasibility of these models in practical applications. In this paper, we propose a novel model to generate high-quality pseudo OOD samples that are akin to IN-Domain (IND) input utterances and thereby improves the performance of OOD detection. To this end, an autoencoder is trained to map an input utterance into a latent code. Moreover, the codes of IND and OOD samples are trained to be indistinguishable by utilizing a generative adversarial network. To provide more supervision signals, an auxiliary classifier is introduced to regularize the generated OOD samples to have indistinguishable intent labels. Experiments show that these pseudo OOD samples generated by our model can be used to effectively improve OOD detection in NLU. Besides, we also demonstrate that the effectiveness of these pseudo OOD data can be further improved by efficiently utilizing unlabeled data.
Yinhe Zheng, Guanyi Chen, Minlie Huang
IEEE ACM Trans. Audio Speech Lang. Process.1
2018 Quantifying Context Overlap for Training Word Embeddings
abstract
Most models for learning word embeddings are trained based on the context information of words, more precisely first order cooccurrence relations.In this paper, a metric is designed to estimate second order cooccurrence relations based on context overlap.The estimated values are further used as the augmented data to enhance the learning of word embeddings by joint training with existing neural word embedding models.Experimental results show that better word vectors can be obtained for word similarity tasks and some downstream NLP tasks by the enhanced approach.
Yimeng Zhuang, Jinghui Xie, Yinhe Zheng
EMNLP3