Jian Sun 0021

dblp:68/4942-21 · DBLP profile ↗
← Back
19ranked-venue papers
0as first author
15since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
YearPublicationVenuePosition
2022 GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy Injection
abstract
Pre-trained models have proved to be powerful in enhancing task-oriented dialog systems. However, current pre-training methods mainly focus on enhancing dialog understanding and generation tasks while neglecting the exploitation of dialog policy. In this paper, we propose GALAXY, a novel pre-trained dialog model that explicitly learns dialog policy from limited labeled dialogs and large-scale unlabeled dialog corpora via semi-supervised learning. Specifically, we introduce a dialog act prediction task for policy optimization during pre-training and employ a consistency regularization term to refine the learned representation with the help of unlabeled dialogs. We also implement a gating mechanism to weigh suitable unlabeled dialog samples. Empirical results show that GALAXY substantially improves the performance of task-oriented dialog systems, and achieves new state-of-the-art results on benchmark datasets: In-Car, MultiWOZ2.0 and MultiWOZ2.1, improving their end-to-end combined scores by 2.5, 5.3 and 5.5 points, respectively. We also show that GALAXY has a stronger few-shot ability than existing models under various low-resource settings. For reproducibility, we release the code and data at https://github.com/siat-nlp/GALAXY.
Wanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu, Zheng Cao 0003, Dermot Liu, Min Yang 0007, Fei Huang 0002, Luo Si, Jian Sun 0021, Yongbin Li 0001
AAAI11
2022 Improving Meta-learning for Low-resource Text Classification and Generation via Memory Imitation
abstract
Yingxiu Zhao, Zhiliang Tian, Huaxiu Yao, Yinhe Zheng, Dongkyu Lee, Yiping Song, Jian Sun, Nevin Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yingxiu Zhao, Zhiliang Tian, Huaxiu Yao, Yinhe Zheng, Yiping Song, Jian Sun 0021, Nevin Lianwen Zhang
ACL (1)7
2022 CGoDial: A Large-Scale Benchmark for Chinese Goal-oriented Dialog Evaluation
abstract
Practical dialog systems need to deal with various knowledge sources, noisy user expressions, and the shortage of annotated data.To better solve the above problems, we propose CGoDial 1 , a new challenging and comprehensive Chinese benchmark for multi-domain Goal-oriented Dialog evaluation.It contains 96,763 dialog sessions, and 574,949 dialog turns totally, covering three datasets with different knowledge sources: 1) a slot-based dialog (SBD) dataset with table-formed knowledge, 2) a flow-based dialog (FBD) dataset with treeformed knowledge, and a retrieval-based dialog (RBD) dataset with candidate-formed knowledge.To bridge the gap between academic benchmarks and spoken dialog scenarios, we either collect data from real conversations or add spoken features to existing datasets via crowdsourcing.The proposed experimental settings include the combinations of training with either the entire training set or a few-shot training set, and testing with either the standard test set or a hard test subset, which can assess model capabilities in terms of general prediction, fast adaptability and reliable robustness.
Yinpei Dai, Wanwei He, Bowen Li 0002, Yuchuan Wu, Zheng Cao 0003, Zhongqi An, Jian Sun 0021, Yongbin Li 0001
EMNLP7
2022 Estimating Soft Labels for Out-of-Domain Intent Detection
abstract
Out-of-Domain (OOD) intent detection is important for practical dialog systems.To alleviate the issue of lacking OOD training samples, some works propose synthesizing pseudo OOD samples and directly assigning one-hot OOD labels to these pseudo samples.However, these one-hot labels introduce noises to the training process because some "hard" pseudo OOD samples may coincide with In-Domain (IND) intents.In this paper, we propose an adaptive soft pseudo labeling (ASoul) method that can estimate soft labels for pseudo OOD samples when training OOD detectors.Semantic connections between pseudo OOD samples and IND intents are captured using an embedding graph.A co-training framework is further introduced to produce resulting soft labels following the smoothness assumption, i.e., close samples are likely to have similar labels.Extensive experiments on three benchmark datasets show that ASoul consistently improves the OOD detection performance and outperforms various competitive baselines.
Hao Lang, Yinhe Zheng, Jian Sun 0021, Fei Huang 0002, Luo Si, Yongbin Li 0001
EMNLP3
2022 Prompt Conditioned VAE: Enhancing Generative Replay for Lifelong Learning in Task-Oriented Dialogue
abstract
Lifelong learning (LL) is vital for advanced task-oriented dialogue (ToD) systems.To address the catastrophic forgetting issue of LL, generative replay methods are widely employed to consolidate past knowledge with generated pseudo samples.However, most existing generative replay methods use only a single taskspecific token to control their models.This scheme is usually not strong enough to constrain the generative model due to insufficient information involved.In this paper, we propose a novel method, prompt conditioned VAE for lifelong learning (PCLL), to enhance generative replay by incorporating tasks' statistics.PCLL captures task-specific distributions with a conditional variational autoencoder, conditioned on natural language prompts to guide the pseudo-sample generation.Moreover, it leverages a distillation process to further consolidate past knowledge by alleviating the noise in pseudo samples.Experiments on natural language understanding tasks of ToD systems demonstrate that PCLL significantly outperforms competitive baselines in building lifelong learning models.We release the code and data at GitHub.
Yingxiu Zhao, Yinhe Zheng, Zhiliang Tian, Jian Sun 0021, Nevin Lianwen Zhang
EMNLP5
2022 A Survey on Neural Open Information Extraction: Current Status and Future Directions
abstract
Open Information Extraction (OpenIE) facilitates domain-independent discovery of relational facts from large corpora. The technique well suits many open-world natural language understanding scenarios, such as automatic knowledge base construction, open-domain question answering, and explicit reasoning. Thanks to the rapid development in deep learning technologies, numerous neural OpenIE architectures have been proposed and achieve considerable performance improvement. In this survey, we provide an extensive overview of the state-of-the-art neural OpenIE models, their key design decisions, strengths and weakness. Then, we discuss limitations of current solutions and the open issues in OpenIE problem itself. Finally we list recent trends that could help expand its scope and applicability, setting up promising directions for future research in OpenIE. To our best knowledge, this paper is the first review on neural OpenIE.
Shaowen Zhou, Bowen Yu 0002, Aixin Sun, Cheng Long 0001, Jian Sun 0021
IJCAI6
2022 Duplex Conversation: Towards Human-like Interaction in Spoken Dialogue Systems
abstract
In this paper, we present Duplex Conversation, a multi-turn, multimodal spoken dialogue system that enables telephone-based agents to interact with customers like a human. We use the concept of full-duplex in telecommunication to demonstrate what a human-like interactive experience should be and how to achieve smooth turn-taking through three subtasks: user state detection, backchannel selection, and barge-in detection. Besides, we propose semi-supervised learning with multimodal data augmentation to leverage unlabeled data to increase model generalization. Experimental results on three sub-tasks show that the proposed method achieves consistent improvements compared with baselines. We deploy the Duplex Conversation to Alibaba intelligent customer service and share lessons learned in production. Online A/B experiments show that the proposed system can significantly reduce response latency by 50%.
Ting-En Lin, Yuchuan Wu, Fei Huang 0002, Luo Si, Jian Sun 0021, Yongbin Li 0001
KDD5
2022 Proton: Probing Schema Linking Information from Pre-trained Language Models for Text-to-SQL Parsing
abstract
The importance of building text-to-SQL parsers which can be applied to new databases has long been acknowledged, and a critical step to achieve this goal is schema linking, i.e., properly recognizing mentions of unseen columns or tables when generating SQLs. In this work, we propose a novel framework to elicit relational structures from large-scale pre-trained language models (PLMs) via a probing procedure based on Poincaré distance metric, and use the induced relations to augment current graph-based parsers for better schema linking. Compared with commonly-used rule-based methods for schema linking, we found that probing relations can robustly capture semantic correspondences, even when surface forms of mentions and entities differ. Moreover, our probing procedure is entirely unsupervised and requires no additional parameters. Extensive experiments show that our framework sets new state-of-the-art performance on three benchmarks. We empirically verify that our probing procedure can indeed find desired relational structures through qualitative analysis.
Bowen Qin, Binyuan Hui, Bowen Li 0002, Min Yang 0007, Bailin Wang, Binhua Li, Jian Sun 0021, Fei Huang 0002, Luo Si, Yongbin Li 0001
KDD8
2022 MMChat: Multi-Modal Chat Dataset on Social Media
abstract
Incorporating multi-modal contexts in conversation is an important step for developing more engaging dialogue systems. In this work, we explore this direction by introducing MMChat: a large scale Chinese multi-modal dialogue corpus (32.4M raw dialogues and 120.84K filtered dialogues). Unlike previous corpora that are crowd-sourced or collected from fictitious movies, MMChat contains image-grounded dialogues collected from real conversations on social media, in which the sparsity issue is observed. Specifically, image-initiated dialogues in common communications may deviate to some non-image-grounded topics as the conversation proceeds. To better investigate this issue, we manually annotate 100K dialogues from MMChat and further filter the corpus accordingly, which yields MMChat-hf. We develop a benchmark model to address the sparsity issue in dialogue generation tasks by adapting the attention routing mechanism on image features. Experiments demonstrate the usefulness of incorporating image features and the effectiveness in handling the sparsity of image features.
Yinhe Zheng, Guanyi Chen, Jian Sun 0021
LREC4
2022 Layout-Aware Information Extraction for Document-Grounded Dialogue: Dataset, Method and Demonstration
abstract
Building document-grounded dialogue systems have received growing interest as documents convey a wealth of human knowledge and commonly exist in enterprises. Wherein, how to comprehend and retrieve information from documents is a challenging research problem. Previous work ignores the visual property of documents and treats them as plain text, resulting in incomplete modality. In this paper, we propose a Layout-aware document-level Information Extraction dataset, LIE, to facilitate the study of extracting both structural and semantic knowledge from visually rich documents (VRDs), so as to generate accurate responses in dialogue systems. LIE contains 62k annotations of three extraction tasks from 4,061 pages in product and official documents, becoming the largest VRD-based information extraction dataset to the best of our knowledge. We also develop benchmark methods that extend the token-based language model to consider layout features like humans. Empirical results show that layout is critical for VRD-based extraction, and system demonstration also verifies that the extracted knowledge can help locate the answers that users care about.
Zhenyu Zhang 0006, Bowen Yu 0002, Haiyang Yu 0003, Tingwen Liu, Cheng Fu 0003, Chengguang Tang, Jian Sun 0021, Yongbin Li 0001
ACM Multimedia8
2022 Unified Dialog Model Pre-training for Task-Oriented Dialog Understanding and Generation
abstract
Recently, pre-training methods have shown remarkable success in task-oriented dialog (TOD) systems. However, most existing pre-trained models for TOD focus on either dialog understanding or dialog generation, but not both. In this paper, we propose SPACE, a novel unified pre-trained dialog model learning from large-scale dialog corpora with limited annotations, which can be effectively fine-tuned on a wide range of downstream dialog tasks. Specifically, SPACE consists of four successive components in a single transformer to maintain a task-flow in TOD systems: (i) a dialog encoding module to encode dialog history, (ii) a dialog understanding module to extract semantic vectors from either user queries or system responses, (iii) a dialog policy module to generate a policy vector that contains high-level semantics of the response, and (iv) a dialog generation module to produce appropriate responses. We design a dedicated pre-training objective for each component. Concretely, we pre-train the dialog encoding module with span mask language modeling to learn contextualized dialog information. To capture the structured dialog semantics, we pre-train the dialog understanding module via a novel tree-induced semi-supervised contrastive learning objective with the help of extra dialog annotations. In addition, we pre-train the dialog policy module by minimizing the ℒ2 distance between its output policy vector and the semantic vector of the response for policy optimization. Finally, the dialog generation model is pre-trained by language modeling. Results show that SPACE achieves state-of-the-art performance on eight downstream dialog benchmarks, including intent prediction, dialog state tracking, and end-to-end dialog modeling. We also show that SPACE has a stronger few-shot ability than existing models under the low-resource setting.
Wanwei He, Yinpei Dai, Min Yang 0007, Jian Sun 0021, Fei Huang 0002, Luo Si, Yongbin Li 0001
SIGIR4
2021 Dynamic Hybrid Relation Exploration Network for Cross-Domain Context-Dependent Semantic Parsing
abstract
Semantic parsing has long been a fundamental problem in natural language processing. Recently, cross-domain context-dependent semantic parsing has become a new focus of research. Central to the problem is the challenge of leveraging contextual information of both natural language queries and database schemas in the interaction history. In this paper, we present a dynamic graph framework that is capable of effectively modelling contextual utterances, tokens, database schemas, and their complicated interaction as the conversation proceeds. The framework employs a dynamic memory decay mechanism that incorporates inductive bias to integrate enriched contextual relation representation, which is further enhanced with a powerful reranking model. At the time of writing, we demonstrate that the proposed framework outperforms all existing models by large margins, achieving new state-of-the-art performance on two large-scale benchmarks, the SParC and CoSQL datasets. Specifically, the model attains a 55.8% question-match and 30.8% interaction-match accuracy on SParC, and a 46.8% question-match and 17.0% interaction-match accuracy on CoSQL.
Binyuan Hui, Ruiying Geng, Qiyu Ren, Binhua Li, Yongbin Li 0001, Jian Sun 0021, Fei Huang 0002, Luo Si, Pengfei Zhu 0001, Xiaodan Zhu 0001
AAAI6
2021 Unsupervised Learning of Deterministic Dialogue Structure with Edge-Enhanced Graph Auto-Encoder
abstract
It is important for task-oriented dialogue systems to discover the dialogue structure (i.e. the general dialogue flow) from dialogue corpora automatically. Previous work models dialogue structure by extracting latent states for each utterance first and then calculating the transition probabilities among states. These two-stage methods ignore the contextual information when calculating the probabilities, which makes the transitions between the states ambiguous. This paper proposes a conversational graph (CG) to represent deterministic dialogue structure where nodes and edges represent the utterance and context information respectively. An unsupervised Edge-Enhanced Graph Auto-Encoder (EGAE) architecture is designed to model local-contextual and global-structural information for conversational graph learning. Furthermore, a self-supervised objective is introduced with the response selection task to guide the unsupervised learning of the dialogue structure. Experimental results on several public datasets demonstrate that the novel model outperforms several alternatives in aggregating utterances with similar semantics. The effectiveness of the learned dialogue structured is also verified by more than 5\% joint accuracy improvement in the downstream task of low resource dialogue state tracking.
Yajing Sun, Yong Shan, Chengguang Tang, Yue Hu 0002, Yinpei Dai, Jing Yu 0007, Jian Sun 0021, Fei Huang 0002, Luo Si
AAAI7
2021 DialogueCSE: Dialogue-based Contrastive Learning of Sentence Embeddings
abstract
Learning sentence embeddings from dialogues has drawn increasing attention due to its low annotation cost and high domain adaptability.Conventional approaches employ the siamese-network for this task, which obtains the sentence embeddings through modeling the context-response semantic relevance by applying a feed-forward network on top of the sentence encoders.However, as the semantic textual similarity is commonly measured through the element-wise distance metrics (e.g.cosine and L2 distance), such architecture yields a large gap between training and evaluating.In this paper, we propose DialogueCSE, a dialogue-based contrastive learning approach to tackle this issue.DialogueCSE first introduces a novel matching-guided embedding (MGE) mechanism, which generates a contextaware embedding for each candidate response embedding (i.e. the context-free embedding) according to the guidance of the multi-turn context-response matching matrices.Then it pairs each context-aware embedding with its corresponding context-free embedding and finally minimizes the contrastive loss across all pairs.We evaluate our model on three multi-turn dialogue datasets: the Microsoft Dialogue Corpus, the Jing Dong Dialogue Corpus, and the E-commerce Dialogue Corpus.Evaluation results show that our approach significantly outperforms the baselines across all three datasets in terms of MAP and Spearman's correlation measures, demonstrating its effectiveness.Further quantitative experiments show that our approach achieves better performance when leveraging more dialogue context and remains robust when less training data is provided.
Rui Wang 0005, Jian Sun 0021, Fei Huang 0002, Luo Si
EMNLP (1)4
2021 Relational Learning with Gated and Attentive Neighbor Aggregator for Few-Shot Knowledge Graph Completion
abstract
Aiming at expanding few-shot relations' coverage in knowledge graphs (KGs), few-shot knowledge graph completion (FKGC) has recently gained more research interests. Some existing models employ a few-shot relation's multi-hop neighbor information to enhance its semantic representation. However, noise neighbor information might be amplified when the neighborhood is excessively sparse and no neighbor is available to represent the few-shot relation. Moreover, modeling and inferring complex relations of one-to-many (1-N), many-to-one (N-1), and many-to-many (N-N) by previous knowledge graph completion approaches requires high model complexity and a large amount of training instances. Thus, inferring complex relations in the few-shot scenario is difficult for FKGC models due to limited training instances. In this paper, we propose a few-shot relational learning with global-local framework to address the above issues. At the global stage, a novel gated and attentive neighbor aggregator is built for accurately integrating the semantics of a few-shot relation's neighborhood, which helps filtering the noise neighbors even if a KG contains extremely sparse neighborhoods. For the local stage, a meta-learning based TransH (MTransH) method is designed to model complex relations and train our model in a few-shot learning fashion. Extensive experiments show that our model outperforms the state-of-the-art FKGC approaches on the frequently-used benchmark datasets NELL-One and Wiki-One. Compared with the strong baseline model MetaR, our model achieves 5-shot FKGC performance improvements of 8.0% on NELL-One and 2.8% on Wiki-One by the metric [email protected]
Guanglin Niu, Yang Li 0218, Chengguang Tang, Ruiying Geng, Hao Wang 0005, Jian Sun 0021, Fei Huang 0002, Luo Si
SIGIR8
2020 Multi-Point Semantic Representation for Intent Classification
abstract
Detecting user intents from utterances is the basis of natural language understanding (NLU) task. To understand the meaning of utterances, some work focuses on fully representing utterances via semantic parsing in which annotation cost is labor-intentsive. While some researchers simply view this as intent classification or frequently asked questions (FAQs) retrieval, they do not leverage the shared utterances among different intents. We propose a simple and novel multi-point semantic representation framework with relatively low annotation cost to leverage the fine-grained factor information, decomposing queries into four factors, i.e., topic, predicate, object/condition, query type. Besides, we propose a compositional intent bi-attention model under multi-task learning with three kinds of attention mechanisms among queries, labels and factors, which jointly combines coarse-grained intent and fine-grained factor information. Extensive experiments show that our framework and model significantly outperform several state-of-the-art approaches with an improvement of 1.35%-2.47% in terms of accuracy.
Jinghan Zhang 0004, Yuxiao Ye, Yue Zhang 0004, Likun Qiu, Yang Li 0218, Zhenglu Yang, Jian Sun 0021
AAAI8
2020 Learning Low-Resource End-To-End Goal-Oriented Dialog for Fast and Reliable System Deployment
abstract
Existing end-to-end dialog systems perform less effectively when data is scarce. To obtain an acceptable success in real-life online services with only a handful of training examples, both fast adaptability and reliable performance are highly desirable for dialog systems. In this paper, we propose the Meta-Dialog System (MDS), which combines the advantages of both meta-learning approaches and human-machine collaboration. We evaluate our methods on a new extended-bAbI dataset and a transformed MultiWOZ dataset for low-resource goal-oriented dialog learning. Experimental results show that MDS significantly outperforms non-meta-learning baselines and can achieve more than 90% per-turn accuracies with only 10 dialogs on the extended-bAbI dataset.
Yinpei Dai, Hangyu Li 0003, Chengguang Tang, Yongbin Li 0001, Jian Sun 0021, Xiaodan Zhu 0001
ACL5
2020 Dynamic Memory Induction Networks for Few-Shot Text Classification
abstract
This paper proposes Dynamic Memory Induction Networks (DMIN) for few-shot text classification.The model utilizes dynamic routing to provide more flexibility to memory-based few-shot learning in order to better adapt the support sets, which is a critical capacity of fewshot classification models.Based on that, we further develop induction models with query information, aiming to enhance the generalization ability of meta-learning.The proposed model achieves new state-of-the-art results on the miniRCV1 and ODIC dataset, improving the best performance (accuracy) by 2∼4%.Detailed analysis is further performed to show the effectiveness of each component.
Ruiying Geng, Binhua Li, Yongbin Li 0001, Jian Sun 0021, Xiaodan Zhu 0001
ACL4
2019 Induction Networks for Few-Shot Text Classification
abstract
Ruiying Geng, Binhua Li, Yongbin Li, Xiaodan Zhu, Ping Jian, Jian Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ruiying Geng, Binhua Li, Yongbin Li 0001, Xiaodan Zhu 0001, Ping Jian, Jian Sun 0021
EMNLP/IJCNLP (1)6