VLDB 2026 Research / reviewers in the wild / expert
Su Zhu
dblp:160/8144
· DBLP profile ↗
36ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0002-9886-9294ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Reducing Tool Hallucination via Reliability AlignmentabstractLarge Language Models (LLMs) have expanded their capabilities beyond language generation to interact with external tools, enabling automation and real-world applications. However, tool hallucinations—where models either select inappropriate tools or misuse them—pose significant challenges, leading to erroneous task execution, increased computational costs, and reduced system reliability. To systematically address this issue, we define and categorize tool hallucinations into two main types: tool selection hallucination and tool usage hallucination. To evaluate and mitigate these issues, we introduce RelyToolBench, which integrates specialized test cases and novel metrics to assess hallucination-aware task success and efficiency. Finally, we propose Relign, a reliability alignment framework that expands the tool-use action space to include indecisive actions, allowing LLMs to defer tool use, seek clarification, or adjust tool selection dynamically. Through extensive experiments, we demonstrate that Relign significantly reduces tool hallucinations, improves task reliability, and enhances the efficiency of LLM tool interactions. The code and data will be publicly available. Hongshen Xu, Su Zhu, Ruisheng Cao, Lu Chen 0002, Kai Yu 0004 |
ICML | 5 |
| 2025 | Unsupervised Text Style Transfer via LLMs and Mask-Filling with Multi-way Interactions
Yuanyuan Liang, Hongshen Xu, Su Zhu, Shuai Fan 0005 |
ICONIP (1) | 5 |
| 2024 | A Birgat Model for Multi-Intent Spoken Language Understanding with Hierarchical Semantic FramesabstractPrevious work on spoken language understanding (SLU) mainly focuses on single-intent settings, where each input utterance merely contains one user intent. This configuration significantly limits the surface form of user utterances and the capacity of output semantics. In this work, we firstly propose a Multi-Intent dataset which is collected from a realistic in-Vehicle dialogue System, called MIVS. The target semantic frame is organized in a 3-layer hierarchical structure to tackle the alignment and assignment problems in multi-intent cases. Accordingly, we devise a BiRGAT model to encode the hierarchy of ontology items, the backbone of which is a dual relational graph attention network. Coupled with the 3-way pointer-generator decoder, our method outperforms traditional sequence labeling and classification-based schemes by a large margin. Ablation study in transfer learning settings further uncovers the poor generalizability of current models in multi-intent cases. Hongshen Xu, Ruisheng Cao, Su Zhu, Hanchong Zhang, Lu Chen 0002, Kai Yu 0004 |
ICASSP | 3 |
| 2024 | Evolving Subnetwork Training for Large Language ModelsabstractLarge language models have ushered in a new era of artificial intelligence research. However, their substantial training costs hinder further development and widespread adoption. In this paper, inspired by the redundancy in the parameters of large language models, we propose a novel training paradigm: Evolving Subnetwork Training (EST). EST samples subnetworks from the layers of the large language model and from commonly used modules within each layer, Multi-Head Attention (MHA) and Multi-Layer Perceptron (MLP). By gradually increasing the size of the subnetworks during the training process, EST can save the cost of training. We apply EST to train GPT2 model and TinyLlama model, resulting in 26.7% FLOPs saving for GPT2 and 25.0% for TinyLlama without an increase in loss on the pre-training dataset. Moreover, EST leads to performance improvements in downstream tasks, indicating that it benefits generalization. Additionally, we provide intuitive theoretical studies based on training dynamics and Dropout theory to ensure the feasibility of EST. Lu Chen 0002, Su Zhu, Kai Yu 0004 |
ICML | 5 |
| 2023 | OPAL: Ontology-Aware Pretrained Language Model for End-to-End Task-Oriented DialogueabstractAbstract This paper presents an ontology-aware pretrained language model (OPAL) for end-to-end task-oriented dialogue (TOD). Unlike chit-chat dialogue models, task-oriented dialogue models fulfill at least two task-specific modules: Dialogue state tracker (DST) and response generator (RG). The dialogue state consists of the domain-slot-value triples, which are regarded as the user’s constraints to search the domain-related databases. The large-scale task-oriented dialogue data with the annotated structured dialogue state usually are inaccessible. It prevents the development of the pretrained language model for the task-oriented dialogue. We propose a simple yet effective pretraining method to alleviate this problem, which consists of two pretraining phases. The first phase is to pretrain on large-scale contextual text data, where the structured information of the text is extracted by the information extracting tool. To bridge the gap between the pretraining method and downstream tasks, we design two pretraining tasks: ontology-like triple recovery and next-text generation, which simulates the DST and RG, respectively. The second phase is to fine-tune the pretrained model on the TOD data. The experimental results show that our proposed method achieves an exciting boost and obtains competitive performance even without any TOD data on CamRest676 and MultiWOZ benchmarks. Zhi Chen 0006, Yuncong Liu, Lu Chen 0002, Su Zhu, Mengyue Wu, Kai Yu 0004 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2022 | UniDU: Towards A Unified Generative Dialogue Understanding FrameworkabstractWith the development of pre-trained language models, remarkable success has been witnessed in dialogue understanding (DU).However, current DU approaches usually employ independent models for each distinct DU task without considering shared knowledge across different DU tasks.In this paper, we propose a unified generative dialogue understanding framework, named UniDU, to achieve effective information exchange across diverse DU tasks.Here, we reformulate all DU tasks into a unified promptbased generative model paradigm.More importantly, a novel model-agnostic multi-task training strategy (MATS) is introduced to dynamically adapt the weights of diverse tasks for best knowledge sharing during training, based on the nature and available data of each task.Experiments on ten DU datasets covering five fundamental DU tasks show that the proposed UniDU framework largely outperforms task-specific well-designed methods on all tasks.MATS also reveals the knowledgesharing structure of these tasks.Finally, UniDU obtains promising performance in the unseen dialogue domain, showing the great potential for generalization. Zhi Chen 0006, Lu Chen 0002, Bei Chen 0008, Libo Qin 0001, Yuncong Liu, Su Zhu, Jian-Guang Lou, Kai Yu 0004 |
SIGDIAL | 6 |
| 2021 | LET: Linguistic Knowledge Enhanced Graph Transformer for Chinese Short Text MatchingabstractChinese short text matching is a fundamental task in natural language processing. Existing approaches usually take Chinese characters or words as input tokens. They have two limitations: 1) Some Chinese words are polysemous, and semantic information is not fully utilized. 2) Some models suffer potential issues caused by word segmentation. Here we introduce HowNet as an external knowledge base and propose a Linguistic knowledge Enhanced graph Transformer (LET) to deal with word ambiguity. Additionally, we adopt the word lattice graph as input to maintain multi-granularity information. Our model is also complementary to pre-trained language models. Experimental results on two Chinese datasets show that our models outperform various typical text matching approaches. Ablation study also indicates that both semantic information and multi-granularity information are important for text matching modeling. Boer Lyu, Lu Chen 0002, Su Zhu, Kai Yu 0004 |
AAAI | 3 |
| 2021 | LGESQL: Line Graph Enhanced Text-to-SQL Model with Mixed Local and Non-Local RelationsabstractRuisheng Cao, Lu Chen, Zhi Chen, Yanbin Zhao, Su Zhu, Kai Yu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ruisheng Cao, Lu Chen 0002, Zhi Chen 0006, Yanbin Zhao, Su Zhu, Kai Yu 0004 |
ACL/IJCNLP (1) | 5 |
| 2021 | ShadowGNN: Graph Projection Neural Network for Text-to-SQL ParserabstractZhi Chen, Lu Chen, Yanbin Zhao, Ruisheng Cao, Zihan Xu, Su Zhu, Kai Yu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhi Chen 0006, Lu Chen 0002, Yanbin Zhao, Ruisheng Cao, Su Zhu, Kai Yu 0004 |
NAACL-HLT | 6 |
| 2021 | Few-Shot NLU with Vector Projection Distance and Abstract Triangular CRF
Su Zhu, Lu Chen 0002, Ruisheng Cao, Zhi Chen 0006, Qingliang Miao, Kai Yu 0004 |
NLPCC (1) | 1 |
| 2020 | Schema-Guided Multi-Domain Dialogue State Tracking with Graph Attention Neural NetworksabstractDialogue state tracking (DST) aims at estimating the current dialogue state given all the preceding conversation. For multi-domain DST, the data sparsity problem is also a major obstacle due to the increased number of state candidates. Existing approaches generally predict the value for each slot independently and do not consider slot relations, which may aggravate the data sparsity problem. In this paper, we propose a Schema-guided multi-domain dialogue State Tracker with graph attention networks (SST) that predicts dialogue states from dialogue utterances and schema graphs which contain slot relations in edges. We also introduce a graph attention matching network to fuse information from utterances and graphs, and a recurrent graph attention network to control state updating. Experiment results show that our approach obtains new state-of-the-art performance on both MultiWOZ 2.0 and MultiWOZ 2.1 benchmarks. Lu Chen 0002, Boer Lv, Su Zhu, Bowen Tan, Kai Yu 0004 |
AAAI | 4 |
| 2020 | Unsupervised Dual Paraphrasing for Two-stage Semantic ParsingabstractOne daunting problem for semantic parsing is the scarcity of annotation.Aiming to reduce nontrivial human labor, we propose a two-stage semantic parsing framework, where the first stage utilizes an unsupervised paraphrase model to convert an unlabeled natural language utterance into the canonical utterance.The downstream naive semantic parser accepts the intermediate output and returns the target logical form.Furthermore, the entire training process is split into two phases: pre-training and cycle learning.Three tailored self-supervised tasks are introduced throughout training to activate the unsupervised paraphrase model.Experimental results on benchmarks OVERNIGHT and GE-OGRANNO demonstrate that our framework is effective and compatible with supervised training. Ruisheng Cao, Su Zhu, Chen Liu 0019, Rao Ma, Yanbin Zhao, Lu Chen 0002, Kai Yu 0004 |
ACL | 2 |
| 2020 | Neural Graph Matching Networks for Chinese Short Text MatchingabstractChinese short text matching usually employs word sequences rather than character sequences to get better performance.However, Chinese word segmentation can be erroneous, ambiguous or inconsistent, which consequently hurts the final matching performance.To address this problem, we propose neural graph matching networks, a novel sentence matching framework capable of dealing with multi-granular input information.Instead of a character sequence or a single word sequence, paired word lattices formed from multiple word segmentation hypotheses are used as input and the model learns a graph representation according to an attentive graph matching mechanism.Experiments on two Chinese datasets show that our models outperform the state-of-the-art short text matching models. Lu Chen 0002, Yanbin Zhao, Boer Lyu, Lesheng Jin, Zhi Chen 0006, Su Zhu, Kai Yu 0004 |
ACL | 6 |
| 2020 | Line Graph Enhanced AMR-to-Text Generation with Mix-Order Graph Attention NetworksabstractEfficient structure encoding for graphs with labeled edges is an important yet challenging point in many graph-based models.This work focuses on AMR-to-text generation -A graph-to-sequence task aiming to recover natural language from Abstract Meaning Representations (AMR).Existing graph-to-sequence approaches generally utilize graph neural networks as their encoders, which have two limitations: 1) The message propagation process in AMR graphs is only guided by the firstorder adjacency information.2) The relationships between labeled edges are not fully considered.In this work, we propose a novel graph encoding framework which can effectively explore the edge relations.We also adopt graph attention networks with higherorder neighborhood information to encode the rich structure in AMR graphs.Experiment results show that our approach obtains new state-of-the-art performance on English AMR benchmark datasets.The ablation analyses also demonstrate that both edge relations and higher-order information are beneficial to graph-to-sequence modeling. Yanbin Zhao, Lu Chen 0002, Zhi Chen 0006, Ruisheng Cao, Su Zhu, Kai Yu 0004 |
ACL | 5 |
| 2020 | A Hierarchical Tracker for Multi-Domain Dialogue State TrackingabstractThe goal of Dialogue State Tracking (DST) is to estimate the current dialogue state given all the preceding conversation. Due to the increased number of state candidates, data sparsity problem is still a major hurdle for multi-domain DST. Existing methods generally choose to predict a value for each possible slot over all domains with quite low efficiency. In this paper, we propose a hierarchical dialogue state tracker which consists of three sequential modules: domain classification, slot detection and value extraction. It predicts domains, slots and values dynamically by given the dialogue history and outputs of the preceding module, which can dramatically improve the model efficiency. Experimental results on MultiWOZ2.1 also show that our approach achieves state-of-the-art joint goal accuracy, and confirm that the hierarchical structure can enhance existing DST models significantly. Su Zhu, Kai Yu 0004 |
ICASSP | 2 |
| 2020 | Jointly Encoding Word Confusion Network and Dialogue Context with BERT for Spoken Language UnderstandingabstractSpoken Language Understanding (SLU) converts hypotheses from automatic speech recognizer (ASR) into structured semantic representations. ASR recognition errors can severely degenerate the performance of the subsequent SLU module. To address this issue, word confusion networks (WCNs) have been used to encode the input for SLU, which contain richer information than 1-best or n-best hypotheses list. To further eliminate ambiguity, the last system act of dialogue context is also utilized as additional input. In this paper, a novel BERT based SLU model (WCN-BERT SLU) is proposed to encode WCNs and the dialogue context jointly. It can integrate both structural information and ASR posterior probabilities of WCNs in the BERT architecture. Experiments on DSTC2, a benchmark of SLU, show that the proposed method is effective and can outperform previous state-of-the-art models significantly. Chen Liu 0019, Su Zhu, Ruisheng Cao, Lu Chen 0002, Kai Yu 0004 |
INTERSPEECH | 2 |
| 2020 | Robust Spoken Language Understanding with RL-Based Value Error Recovery
Chen Liu 0019, Su Zhu, Lu Chen 0002, Kai Yu 0004 |
NLPCC (1) | 2 |
| 2020 | Memory Attention Neural Network for Multi-domain Dialogue State Tracking
Zhi Chen 0006, Lu Chen 0002, Su Zhu, Kai Yu 0004 |
NLPCC (1) | 4 |
| 2020 | Dual Learning for Semi-Supervised Natural Language UnderstandingabstractNatural language understanding (NLU) converts sentences into structured semantic forms. The paucity of annotated training samples is still a fundamental challenge of NLU. To solve this data sparsity problem, previous work based on semi-supervised learning mainly focuses on exploiting unlabeled sentences. In this work, we introduce a dual task of NLU, semantic-to-sentence generation (SSG), and propose a new framework for semi-supervised NLU with the corresponding dual model. The framework is composed of dual pseudo-labeling and dual learning method, which enables an NLU model to make full use of data (labeled and unlabeled) through a closed-loop of the primal and dual tasks. By incorporating the dual task, the framework can exploit pure semantic forms as well as unlabeled sentences, and further improve the NLU and SSG models iteratively in the closed-loop. The proposed approaches are evaluated on two public datasets (ATIS and SNIPS). Experiments in the semi-supervised setting show that our methods can outperform various baselines significantly, and extensive ablation studies are conducted to verify the effectiveness of our framework. Finally, our method can also achieve the state-of-the-art performance on the two datasets in the supervised setting. Su Zhu, Ruisheng Cao, Kai Yu 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Prior Knowledge Driven Label Embedding for Slot Filling in Natural Language UnderstandingabstractTraditional slot filling in natural language understanding (NLU) predicts a one-hot vector for each word. This form of label representation lacks semantic correlation modeling, which leads to severe data sparsity problem, especially when adapting an NLU model to a new domain. To address this issue, a novel label embedding based slot filling framework is proposed in this article. Here, distributed label embedding is constructed for each slot using prior knowledge. Three encoding methods are investigated to incorporate different kinds of prior knowledge about slots: atomic concepts, slot descriptions, and slot exemplars. The proposed label embeddings tend to share text patterns and reuses data with different slot labels. This makes it useful for adaptive NLU with limited data. Also, since label embedding is independent of NLU model, it is compatible with almost all deep learning based slot filling models. The proposed approaches are evaluated on three datasets. Experiments on single domain and domain adaptation tasks show that label embedding achieves significant performance improvement over traditional one-hot label representation as well as advanced zero-shot approaches. Su Zhu, Rao Ma, Kai Yu 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Semantic Parsing with Dual LearningabstractSemantic parsing converts natural language queries into structured logical forms.The paucity of annotated training samples is a fundamental challenge in this field.In this work, we develop a semantic parsing framework with the dual learning algorithm, which enables a semantic parser to make full use of data (labeled and even unlabeled) through a dual-learning game.This game between a primal model (semantic parsing) and a dual model (logical form to query) forces them to regularize each other, and can achieve feedback signals from some prior-knowledge.By utilizing the prior-knowledge of logical form structures, we propose a novel reward signal at the surface and semantic levels which tends to generate complete and reasonable logical forms.Experimental results show that our approach achieves new state-of-the-art performance on ATIS dataset and gets competitive performance on OVERNIGHT dataset. Ruisheng Cao, Su Zhu, Chen Liu 0019, Kai Yu 0004 |
ACL (1) | 2 |
| 2019 | Data Augmentation with Atomic Templates for Spoken Language UnderstandingabstractZijian Zhao, Su Zhu, Kai Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Su Zhu, Kai Yu 0004 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | A Hierarchical Decoding Model for Spoken Language Understanding from Unaligned DataabstractSpoken language understanding (SLU) systems can be trained on two types of labelled data: aligned or unaligned. Unaligned data do not require word by word annotation and is easier to be obtained. In the paper, we focus on spoken language understanding from unaligned data whose annotation is a set of act-slot-value triples. Previous works usually focus on improve slot-value pair prediction and estimate dialogue act types separately, which ignores the hierarchical structure of the act-slot-value triples. Here, we propose a novel hierarchical decoding model which dynamically parses act, slot and value in a structured way and employs pointer network to handle out-of-vocabulary (OOV) values. Experiments on DSTC2 dataset, a benchmark unaligned dataset, show that the proposed model not only outperforms previous state-of-the-art model, but also can be generalized effectively and efficiently to unseen act-slot type pairs and OOV values. Su Zhu, Kai Yu 0004 |
ICASSP | 2 |
| 2019 | Robust Spoken Language Understanding with Acoustic and Domain KnowledgeabstractSpoken language understanding (SLU) converts user utterances into structured semantic forms. There are still two main issues for SLU: robustness to ASR-errors and the data sparsity of new and extended domains. In this paper, we propose a robust SLU system by leveraging both acoustic and domain knowledge. We extract audio features by training ASR models on a large number of utterances without semantic annotations. For exploiting domain knowledge, we design lexicon features from the domain ontology and propose an error elimination algorithm to help predicted values recovered from ASR-errors. The results of CATSLU challenge show that our systems can outperform all of the other teams across four domains. Chen Liu 0019, Su Zhu, Kai Yu 0004 |
ICMI | 3 |
| 2019 | CATSLU: The 1st Chinese Audio-Textual Spoken Language Understanding ChallengeabstractSpoken language understanding (SLU) is a key component of conversational dialogue systems, which converts user utterances into semantic representations. The previous works almost focus on parsing semantic from textual inputs (top hypothesis of speech recognition and even manual transcripts) while losing information hidden in the audio. We herein describe the 1st Chinese Audio-Textual Spoken Language Understanding Challenge (CATSLU) which introduces a new dataset with audio-textual information, multiple domains and domain knowledge. We introduce two scenarios of audio-textual SLU in which participants are encouraged to utilize data of other domains or not. In this paper, we will describe the challenge and results. Su Zhu, Tiejun Zhao, Chengqing Zong, Kai Yu 0004 |
ICMI | 1 |
| 2018 | Semi-Supervised Training Using Adversarial Multi-Task Learning for Spoken Language UnderstandingabstractSpoken language understanding (SLU) usually requires human semantic annotation on collected data, but the process is expensive. In order to make better use of unlabeled data for robust SLU, we propose an adversarial multi-task learning method by merging a bidirectional language model (BLM) and a slot tagging model (STM). As a secondary objective, the BLM is used to learn generalized and unsupervised knowledge with abundant unlabeled data and improve the performance of STM on unseen data samples. We construct a shared space for both tasks and independent private spaces for each task respectively. Additional adversarial task discriminator is also used to obtain more task - independent sharing information. Experiments show that the proposed approaches achieve the state-of-the-art performance on the small scale ATIS benchmark and significantly improve the semi -supervised performance on a large-scale dataset. Ouyu Lan, Su Zhu, Kai Yu 0004 |
ICASSP | 2 |
| 2018 | Robust Spoken Language Understanding with Unsupervised ASR-Error AdaptationabstractRobustness to errors produced by automatic speech recognition (ASR) is essential for Spoken Language Understanding (SLU). Traditional robust SLU typically needs ASR hypotheses with semantic annotations for training. However, semantic annotation is very expensive, and the corresponding ASR system may change frequently. Here, we propose a novel unsupervised ASR-error adaptation method, obviating the need of annotated ASR hypotheses. It only requires semantically annotated transcripts for the slot-tagging task and the transcripts paired with hypotheses for an input sentence reconstruction task. In this method, feature encoders which share part of the parameters are exploited to enforce the tasks in a similar feature space. Therefore, the transcript side slot-tagging model can be transferred to ASR hypotheses side easily. Experiments show that the proposed approach can yield significant improvement over strong baselines, and achieve performance very close to the oracle system. Su Zhu, Ouyu Lan, Kai Yu 0004 |
ICASSP | 1 |
| 2018 | Concept Transfer Learning for Adaptive Language UnderstandingabstractConcept definition is important in language understanding (LU) adaptation since literal definition difference can easily lead to data sparsity even if different data sets are actually semantically correlated.To address this issue, in this paper, a novel concept transfer learning approach is proposed.Here, substructures within literal concept definition are investigated to reveal the relationship between concepts.A hierarchical semantic representation for concepts is proposed, where a semantic slot is represented as a composition of atomic concepts.Based on this new hierarchical representation, transfer learning approaches are developed for adaptive LU.The approaches are applied to two tasks: value set mismatch and domain adaptation, and evaluated on two LU benchmarks: ATIS and DSTC 2&3.Thorough empirical studies validate both the efficiency and effectiveness of the proposed method.In particular, we achieve state-ofthe-art performance (F 1 -score 96.08%) on ATIS by only using lexicon features. Su Zhu, Kai Yu 0004 |
SIGDIAL Conference | 1 |
| 2017 | Encoder-decoder with focus-mechanism for sequence labelling based spoken language understandingabstractThis paper investigates the framework of encoder-decoder with attention for sequence labelling based spoken language understanding. We introduce Bidirectional Long Short Term Memory - Long Short Term Memory networks (BLSTM-LSTM) as the encoder-decoder model to fully utilize the power of deep learning. In the sequence labelling task, the input and output sequences are aligned word by word, while the attention mechanism cannot provide the exact alignment. To address this limitation, we propose a novel focus mechanism for encoder-decoder framework. Experiments on the standard ATIS dataset showed that BLSTM-LSTM with focus mechanism defined the new state-of-the-art by outperforming standard BLSTM and attention based encoder-decoder. Further experiments also show that the proposed model is more robust to speech recognition errors. Su Zhu, Kai Yu 0004 |
ICASSP | 1 |
| 2016 | Hybrid Dialogue State Tracking for Real World Human-to-Human Dialogues
Su Zhu, Lu Chen 0002, Siqiu Yao, Xueyang Wu 0001, Kai Yu 0004 |
INTERSPEECH | 2 |
| 2016 | Evolvable dialogue state tracking for statistical dialogue management
Kai Yu 0004, Lu Chen 0002, Qizhe Xie, Su Zhu |
Frontiers Comput. Sci. | 5 |
| 2015 | Recurrent Polynomial Network for Dialogue State Tracking with Mismatched Semantic ParsersabstractRecently, constrained Markov Bayesian polynomial (CMBP) has been proposed as a data-driven rule-based model for dialog state tracking (DST).CMBP is an approach to bridge rule-based models and statistical models.Recurrent Polynomial Network (RPN) is a recent statistical framework taking advantages of rulebased models and can achieve state-ofthe-art performance on the data corpora of DSTC-3, outperforming all submitted trackers in DSTC-3 including RNN.It is widely acknowledged that SLU's reliability influences tracker's performance greatly, especially in cases where the training SLU is poorly matched to the testing SLU.In this paper, this effect is analyzed in detail for RPN.Experiments show that RPN's tracking result is consistently the best compared to rule-based and statistical models investigated on different SLUs including mismatched ones and demonstrate RPN's is very robust to mismatched semantic parsers. Qizhe Xie, Su Zhu, Lu Chen 0002, Kai Yu 0004 |
SIGDIAL Conference | 3 |
| 2015 | Constrained Markov Bayesian Polynomial for Efficient Dialogue State TrackingabstractDialogue state tracking (DST) is a process to estimate the distribution of the dialogue states at each dialogue turn given the interaction history. Although data-driven statistical approaches are of most interest, there have been attempts of using rule-based methods for DST, due to their simplicity, efficiency and portability. However, the performance of these methods are usually not competitive to data-driven tracking approaches and it is not possible to improve the DST performance when training data are available. In this paper, a novel hybrid framework, constrained Markov Bayesian polynomial (CMBP), is proposed to formulate rule-based DST in a general way and allow data-driven rule generation. Here, a DST rule is defined as a polynomial function of a set of probabilities satisfying certain linear constraints. Prior knowledge is encoded in these constraints. Under reasonable assumptions, CMBP optimization can be converted to a constrained integer linear programming problem. The integer coefficient CMBP model is further extended to CMBP with real coefficients by applying grid search. CMBP was evaluated on the data corpora of the first, the second, and the third Dialog State Tracking Challenge (DSTC-1/2/3). Experiments showed that CMBP has good generalization ability and can significantly outperform both traditional rule-based approaches and data-driven statistical approaches with similar feature set. Compared with the state-of-the-art statistical DST approaches with much richer features, CMBP is also competitive. Kai Yu 0004, Lu Chen 0002, Su Zhu |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | The SJTU System for Dialog State Tracking Challenge 2abstractDialog state tracking challenge provides a common testbed for state tracking al-gorithms. This paper describes the SJTU system submitted to the second Dialogue State Tracking Challenge in detail. In the system, a statistical semantic parser is used to generate refined semantic hypothe-ses. A large number of features are then derived based on the semantic hypothe-ses and the dialogue log information. The final tracker is a combination of a rule-based model, a maximum entropy and a deep neural network model. The SJTU system significantly outperformed all the baselines and showed competitive perfor-mance in DSTC 2. 1 Lu Chen 0002, Su Zhu, Kai Yu 0004 |
SIGDIAL Conference | 3 |
| 2014 | A generalized rule based tracker for dialogue state trackingabstractDialogue state tracking plays an important role in statistical dialogue management. Domain-independent rule-based approaches are attractive due to their efficiency, portability and interpretability. However, recent rule-based models are still not quite competitive to statistical tracking approaches. In this paper, a novel framework is proposed to formulate rule-based models in a general way. In the framework, a rule is considered as a special kind of polynomial function satisfying certain linear constraints. Under some particular definitions and assumptions, rule-based models can be seen as feasible solutions of an integer linear programming problem. Experiments showed that the proposed approach can not only achieve competitive performance compared to statistical approaches, but also have good generalisation ability. It is one of the only two entries that outperformed all the four baselines in the third Dialog State Tracking Challenge. Lu Chen 0002, Su Zhu, Kai Yu 0004 |
SLT | 3 |
| 2014 | Semantic parser enhancement for dialogue domain extension with little dataabstractStatistical semantic parser trained on sufficient in-domain data has shown robustness to speech recognition errors in end-to-end spoken dialogue systems. However, when the dialogue domain is extended, due to the introduction of new semantic slots, values and unknown speech pattern, the parsing performance may significantly degrade. Effective re-training of statistical semantic parser is therefore important. This paper describes a novel semantic parser enhancement approach for domain extension with very little new data. It employs automatic pseudo-data generation for parser re-training and domain independent rescoring to further improve parsing performance. The approach was evaluated on the DSTC3 (the third Dialog State Tracking Challenge) data corpus. Experiments showed that the proposed approach can yield consistent and significant improvements across all metrics of semantic parsing and dialog state tracking. Su Zhu, Lu Chen 0002, Kai Yu 0004 |
SLT | 1 |