EDBT 2026 Demo / reviewers in the wild / expert
Lu Chen 0002
dblp:69/157-2
· DBLP profile ↗
64ranked-venue papers
7as first author
37since 2021 · last 2025
0000-0001-8687-4806ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 6 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NeuSym-RAG: Hybrid Neural Symbolic Retrieval with Multiview Structuring for PDF Question AnsweringabstractThe increasing number of academic papers poses significant challenges for researchers to efficiently acquire key details. While retrieval augmented generation (RAG) shows great promise in large language model (LLM) based automated question answering, previous works often isolate neural and symbolic retrieval despite their complementary strengths. Moreover, conventional single-view chunking neglects the rich structure and layout of PDFs, e.g., sections and tables. In this work, we propose NeuSym-RAG, a hybrid neural symbolic retrieval framework which combines both paradigms in an interactive process. By leveraging multi-view chunking and schema-based parsing, NeuSym-RAG organizes semi-structured PDF content into both the relational database and vectorstore, enabling LLM agents to iteratively gather context until sufficient to generate answers. Experiments on three full PDF-based QA datasets, including a self-annotated one AirQA-Real, show that NeuSym-RAG stably defeats both the vector-based RAG and various structured baselines, highlighting its capacity to unify both retrieval schemes and utilize multiple views. Ruisheng Cao, Hanchong Zhang, Tiancheng Huang 0001, Zhangyi Kang, Liangtai Sun, Yuxun Miao, Shuai Fan 0005, Lu Chen 0002, Kai Yu 0004 |
ACL (1) | 10 |
| 2025 | From Generalist to Specialist: A Survey of Large Language Models for ChemistryabstractLarge Language Models (LLMs) have significantly transformed our daily life and established a new paradigm in natural language processing (NLP). However, the predominant pretraining of LLMs on extensive web-based texts remains insufficient for advanced scientific discovery, particularly in chemistry. The scarcity of specialized chemistry data, coupled with the complexity of multi-modal data such as 2D graph, 3D structure and spectrum, present distinct challenges. Although several studies have reviewed Pretrained Language Models (PLMs) in chemistry, there is a conspicuous absence of a systematic survey specifically focused on chemistry-oriented LLMs. In this paper, we outline methodologies for incorporating domain-specific chemistry knowledge and multi-modal information into LLMs, we also conceptualize chemistry LLMs as agents using chemistry tools and investigate their potential to accelerate scientific research. Additionally, we conclude the existing benchmarks to evaluate chemistry ability of LLMs. Finally, we critically examine the current challenges and identify promising directions for future research. Through this comprehensive survey, we aim to assist researchers in staying at the forefront of developments in chemistry LLMs and to inspire innovative applications in the field. Yang Han 0007, Ziping Wan, Lu Chen 0002, Kai Yu 0004 |
COLING | 3 |
| 2025 | ChatCite: LLM Agent with Human Workflow Guidance for Comparative Literature SummaryabstractThe literature review is an indispensable step in the research process. It provides the benefit of comprehending the research problem and understanding the current research situation while conducting a comparative analysis of prior works. However, literature summary is challenging and time consuming. The previous LLM-based studies on literature review mainly focused on the complete process, including literature retrieval, screening, and summarization. However, for the summarization step, simple CoT method often lacks the ability to provide extensive comparative summary. In this work, we firstly focus on the independent literature summarization step and introduce ChatCite, an LLM agent with human workflow guidance for comparative literature summary. This agent, by mimicking the human workflow, first extracts key elements from relevant literature and then generates summaries using a Reflective Incremental Mechanism. In order to better evaluate the quality of the generated summaries, we devised a LLM-based automatic evaluation metric, G-Score, in refer to the human evaluation criteria. The ChatCite agent outperformed other models in various dimensions in the experiments. The literature summaries generated by ChatCite can also be directly used for drafting literature reviews. Lu Chen 0002, Aiwei Liu, Kai Yu 0004, Lijie Wen 0001 |
COLING | 2 |
| 2025 | Converging to a Lingua Franca: Evolution of Linguistic Regions and Semantics Alignment in Multilingual Large Language ModelsabstractLarge language models (LLMs) have demonstrated remarkable performance, particularly in multilingual contexts. While recent studies suggest that LLMs can transfer skills learned in one language to others, the internal mechanisms behind this ability remain unclear. We observed that the neuron activation patterns of LLMs exhibit similarities when processing the same language, revealing the existence and location of key linguistic regions. Additionally, we found that neuron activation patterns are similar when processing sentences with the same semantic meaning in different languages. This indicates that LLMs map semantically identical inputs from different languages into a “Lingua Franca”, a common semantic latent space that allows for consistent processing across languages. This semantic alignment becomes more pronounced with training and increased model size, resulting in a more language-agnostic activation pattern. Moreover, we found that key linguistic neurons are concentrated in the first and last layers of LLMs, becoming denser in the first layers as training progresses. Experiments on BLOOM and LLaMA2 support these findings, highlighting the structural evolution of multilingual LLMs during training and scaling up. This paper provides insights into the internal workings of LLMs, offering a foundation for future improvements in their cross-lingual capabilities. Hongchuan Zeng, Senyu Han, Lu Chen 0002, Kai Yu 0004 |
COLING | 3 |
| 2025 | Alignment for Efficient Tool Calling of Large Language ModelsabstractRecent advancements in tool learning have enabled large language models (LLMs) to integrate external tools, enhancing their task performance by expanding their knowledge boundaries.However, relying on tools often introduces trade-offs between performance, speed, and cost, with LLMs sometimes exhibiting overreliance and overconfidence in tool usage.This paper addresses the challenge of aligning LLMs with their knowledge boundaries to make more intelligent decisions about tool invocation.We propose a multi-objective alignment framework that combines probabilistic knowledge boundary estimation with dynamic decision-making, allowing LLMs to better assess when to invoke tools based on their confidence.Our framework includes two methods for knowledge boundary estimation-consistency-based and absolute estimation-and two training strategies for integrating these estimates into the model's decision-making process.Experimental results on various tool invocation scenarios demonstrate the effectiveness of our framework, showing significant improvements in tool efficiency by reducing unnecessary tool usage. Hongshen Xu, Shuai Fan 0005, Lu Chen 0002, Kai Yu 0004 |
EMNLP | 7 |
| 2025 | Heads up! Large Language Models Can Perform Tasks Without Your Instruction via Selective Attention Head MaskingabstractLarge language models (LLMs) consist of numerous Transformer modules, and while the models can perform various functions, it remains an open question of how these modules are combined to elicit distinct inherent functionalities. In this paper, we investigate the modules inside LLMs and demonstrate that, by simply masking or retaining specific attention heads during inference, LLMs can exhibit specific task functionalities without requiring explicit instructions or modifications to the model parameters. Experiments across various models and tasks reveal that LLMs inherently encode “functional pathways”, the structured groups of interdependent attention heads that are crucial for executing specific tasks. These pathways not only govern the model’s functional behaviors but also enhance parameter efficiency, as suppressing attention heads outside the pathway can improve task performance. The code is available in this repository: https://github.com/OpenDFM/HeadsUp. Senyu Han, Hongchuan Zeng, Kai Yu 0004, Lu Chen 0002 |
ICML | 4 |
| 2025 | Reducing Tool Hallucination via Reliability AlignmentabstractLarge Language Models (LLMs) have expanded their capabilities beyond language generation to interact with external tools, enabling automation and real-world applications. However, tool hallucinations—where models either select inappropriate tools or misuse them—pose significant challenges, leading to erroneous task execution, increased computational costs, and reduced system reliability. To systematically address this issue, we define and categorize tool hallucinations into two main types: tool selection hallucination and tool usage hallucination. To evaluate and mitigate these issues, we introduce RelyToolBench, which integrates specialized test cases and novel metrics to assess hallucination-aware task success and efficiency. Finally, we propose Relign, a reliability alignment framework that expands the tool-use action space to include indecisive actions, allowing LLMs to defer tool use, seek clarification, or adjust tool selection dynamically. Through extensive experiments, we demonstrate that Relign significantly reduces tool hallucinations, improves task reliability, and enhances the efficiency of LLM tool interactions. The code and data will be publicly available. Hongshen Xu, Su Zhu, Ruisheng Cao, Lu Chen 0002, Kai Yu 0004 |
ICML | 8 |
| 2025 | MS-BART: Unified Modeling of Mass Spectra and Molecules for Structure ElucidationabstractMass spectrometry (MS) plays a critical role in molecular identification, significantly advancing scientific discovery. However, structure elucidation from MS data remains challenging due to the scarcity of annotated spectra. While large-scale pretraining has proven effective in addressing data scarcity in other domains, applying this paradigm to mass spectrometry is hindered by the complexity and heterogeneity of raw spectral signals.
To address this, we propose MS-BART, a unified modeling framework that maps mass spectra and molecular structures into a shared token vocabulary, enabling cross-modal learning through large-scale pretraining on reliably computed fingerprint–molecule datasets. Multi-task pretraining objectives further enhance MS-BART's generalization by jointly optimizing denoising and translation task. The pretrained model is subsequently transferred to experimental spectra through finetuning on fingerprint predictions generated with MIST, a pre-trained spectral inference model, thereby enhancing robustness to real-world spectral variability. While finetuning alleviates the distributional difference, MS-BART still suffers molecular hallucination and requires further alignment. We therefore introduce a chemical feedback mechanism that guides the model toward generating molecules closer to the reference structure.
Extensive evaluations demonstrate that MS-BART achieves SOTA performance across 5/12 key metrics on MassSpecGym and NPLIB1 and is faster by one order of magnitude than competing diffusion-based methods, while comprehensive ablation studies systematically validate the model's effectiveness and robustness.
We provide the data and code at [https://github.com/OpenDFM/MS-BART](https://github.com/OpenDFM/MS-BART). Yang Han 0007, Kai Yu 0004, Lu Chen 0002 |
NeurIPS | 5 |
| 2025 | Task-Specific Data Selection for Instruction Tuning via Monosemantic Neuronal ActivationsabstractInstruction tuning improves the ability of large language models (LLMs) to follow diverse human instructions, but achieving strong performance on specific target tasks remains challenging. A critical bottleneck is selecting the most relevant data to maximize task-specific performance. Existing data selection approaches include unstable influence-based methods and more stable distribution alignment methods, the latter of which critically rely on the underlying sample representation. In practice, most distribution alignment methods, from shallow features (e.g., BM25) to neural embeddings (e.g., BGE, LLM2Vec), may fail to capture how the model internally processes samples. To bridge this gap, we adopt a model-centric strategy in which each sample is represented by its neuronal activation pattern in the model, directly reflecting internal computation. However, directly using raw neuron activations leads to spurious similarity between unrelated samples due to neuron polysemanticity, where a single neuron may respond to multiple, unrelated concepts. To address this, we employ sparse autoencoders to disentangle polysemantic activations into sparse, monosemantic representations, and introduce a dedicated similarity metric for this space to better identify task-relevant data. Comprehensive experiments across multiple instruction datasets, models, tasks, and selection ratios show that our approach consistently outperforms existing data selection baselines in both stability and task-specific performance. Gonghu Shang, Zhi Chen 0006, Libo Qin 0001, Yijie Luo, Hongshen Xu, Shuai Fan 0005, Kai Yu 0004, Lu Chen 0002 |
NeurIPS | 10 |
| 2025 | MULTI: multimodal understanding leaderboard with text and images
Lu Chen 0002, Jingkai Yang, Yichuan Ma, Hailin Wen, Jinyu Cai, Yingzi Ma, Situo Zhang, Zihan Zhao 0001, Liangtai Sun, Kai Yu 0004 |
Sci. China Inf. Sci. | 3 |
| 2024 | SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific ResearchabstractRecently, there has been growing interest in using Large Language Models (LLMs) for scientific research. Numerous benchmarks have been proposed to evaluate the ability of LLMs for scientific research. However, current benchmarks are mostly based on pre-collected objective questions. This design suffers from data leakage problem and lacks the evaluation of subjective Q/A ability. In this paper, we propose SciEval, a comprehensive and multi-disciplinary evaluation benchmark to address these issues. Based on Bloom's taxonomy, SciEval covers four dimensions to systematically evaluate scientific research ability. In particular, we design a "dynamic" subset based on scientific principles to prevent evaluation from potential data leakage. Both objective and subjective questions are included in SciEval. These characteristics make SciEval a more effective benchmark for scientific research ability evaluation of LLMs. Comprehensive experiments on most advanced LLMs show that, although GPT-4 achieves SOTA performance compared to other LLMs, there is still substantial room for improvement, especially for dynamic questions. The codes and data are publicly available on https://github.com/OpenDFM/SciEval. Liangtai Sun, Yang Han 0007, Zihan Zhao 0001, Zhennan Shen, Baocai Chen, Lu Chen 0002, Kai Yu 0004 |
AAAI | 7 |
| 2024 | IBSEN: Director-Actor Agent Collaboration for Controllable and Interactive Drama Script GenerationabstractLarge language models have demonstrated their capabilities in storyline creation and humanlike character role-playing.Current language model agents mainly focus on reasonable behaviors from the level of individuals, and their behaviors might be hard to constraint on the level of the whole storyline.In this paper we introduce IBSEN, a director-actor coordinate agent framework that generates drama scripts and makes the plot played by agents more controllable.The director agent writes plot outlines that the user desires to see, instructs the actor agents to role-play their characters, and reschedules the plot when human players participate in the scenario to ensure the plot is progressing towards the objective.To evaluate the framework, we create a novel drama plot that involves several actor agents and check the interactions between them under the instruction of the director agent.Evaluation results show that our framework could generate complete, diverse drama scripts from only a rough outline of plot objectives, meanwhile maintaining the characteristics of characters in the drama.Our codes and prompts are available at https://github.com/OpenDFM/ibsen. Senyu Han, Lu Chen 0002, Li-Min Lin, Zhengshan Xu, Kai Yu 0004 |
ACL (1) | 2 |
| 2024 | Multilingual Brain Surgeon: Large Language Models Can Be Compressed Leaving No Language behindabstractLarge Language Models (LLMs) have ushered in a new era in Natural Language Processing, but their massive size demands effective compression techniques for practicality. Although numerous model compression techniques have been investigated, they typically rely on a calibration set that overlooks the multilingual context and results in significant accuracy degradation for low-resource languages. This paper introduces Multilingual Brain Surgeon (MBS), a novel calibration data sampling method for multilingual LLMs compression. MBS overcomes the English-centric limitations of existing methods by sampling calibration data from various languages proportionally to the language distribution of the model training datasets. Our experiments, conducted on the BLOOM multilingual LLM, demonstrate that MBS improves the performance of existing English-centric compression methods, especially for low-resource languages. We also uncover the dynamics of language interaction during compression, revealing that the larger the proportion of a language in the training set and the more similar the language is to the calibration language, the better performance the language retains after compression. In conclusion, MBS presents an innovative approach to compressing multilingual LLMs, addressing the performance disparities and improving the language inclusivity of existing compression techniques. Keywords: Large Language Model, Multilingual Model Compression Hongchuan Zeng, Hongshen Xu, Lu Chen 0002, Kai Yu 0004 |
LREC/COLING | 3 |
| 2024 | Is LLM a Reliable Reviewer? A Comprehensive Evaluation of LLM on Automatic Paper Reviewing TasksabstractThe use of large language models (LLM), especially ChatGPT, to help with research has come into practice. Researchers use it for timely advice and hope to obtain in-depth feedback. However, can LLM be a qualified and reliable reviewer? Although there already exist several review-related datasets, few works have carefully and thoroughly inspected model’s capability as a reviewer, especially the correctness of generated reviews. In this paper, we first evaluate GPT-3.5 and GPT-4 (the current top-performing LLM) on 2 types of tasks under different settings: the score prediction task and the review generation task. In addition, we propose a dataset containing 197 review-revision multiple-choice questions (RR-MCQ) with detailed labels from the review-rebuttal forum in ICLR-2023. By asking questions from technical details to the overall presentation and quality, our RR-MCQ data provides a more complete model ability assessment. The results show that LLM is generally helpful, but great caution is needed as it always makes mistakes. Although it can give passable decisions (> 60% accuracy) on single options, completely correct answers are still rare (about 20%); models are still weak on long paper processing, zero-shot scoring, and giving critical feedback like human reviewers. Ruiyang Zhou, Lu Chen 0002, Kai Yu 0004 |
LREC/COLING | 2 |
| 2024 | Label-Aware Auxiliary Learning for Dialogue State TrackingabstractDialogue State Tracking (DST) is an essential part of task-oriented dialogue systems. Many existing methods try to utilize external dialogue datasets to improve the performance of DST models. Instead of previous methods, in this paper, we propose Label-Aware Auxiliary Learning for DST (LAL-DST) which focuses on exploiting the abundant internal information of the target DST dataset to improve the performance. We design label-aware auxiliary tasks, in which we apply noising functions to either the dialogue history or the belief state label and take the concatenation of them as input. The goal of each task is to restore the corrupted context. During the training process, we first further train the large pre-trained language model on the auxiliary tasks, then fine-tune it on DST. Through the experimental results, we empirically show the effect of LAL-DST by the performance improvements it brings to MultiWOZ2.0 and WOZ. Yuncong Liu, Lu Chen 0002, Kai Yu 0004 |
ICASSP | 2 |
| 2024 | A Birgat Model for Multi-Intent Spoken Language Understanding with Hierarchical Semantic FramesabstractPrevious work on spoken language understanding (SLU) mainly focuses on single-intent settings, where each input utterance merely contains one user intent. This configuration significantly limits the surface form of user utterances and the capacity of output semantics. In this work, we firstly propose a Multi-Intent dataset which is collected from a realistic in-Vehicle dialogue System, called MIVS. The target semantic frame is organized in a 3-layer hierarchical structure to tackle the alignment and assignment problems in multi-intent cases. Accordingly, we devise a BiRGAT model to encode the hierarchy of ontology items, the backbone of which is a dual relational graph attention network. Coupled with the 3-way pointer-generator decoder, our method outperforms traditional sequence labeling and classification-based schemes by a large margin. Ablation study in transfer learning settings further uncovers the poor generalizability of current models in multi-intent cases. Hongshen Xu, Ruisheng Cao, Su Zhu, Hanchong Zhang, Lu Chen 0002, Kai Yu 0004 |
ICASSP | 6 |
| 2024 | Evolving Subnetwork Training for Large Language ModelsabstractLarge language models have ushered in a new era of artificial intelligence research. However, their substantial training costs hinder further development and widespread adoption. In this paper, inspired by the redundancy in the parameters of large language models, we propose a novel training paradigm: Evolving Subnetwork Training (EST). EST samples subnetworks from the layers of the large language model and from commonly used modules within each layer, Multi-Head Attention (MHA) and Multi-Layer Perceptron (MLP). By gradually increasing the size of the subnetworks during the training process, EST can save the cost of training. We apply EST to train GPT2 model and TinyLlama model, resulting in 26.7% FLOPs saving for GPT2 and 25.0% for TinyLlama without an increase in loss on the pre-training dataset. Moreover, EST leads to performance improvements in downstream tasks, indicating that it benefits generalization. Additionally, we provide intuitive theoretical studies based on training dynamics and Dropout theory to ensure the feasibility of EST. Lu Chen 0002, Su Zhu, Kai Yu 0004 |
ICML | 2 |
| 2024 | CoE-SQL: In-Context Learning for Multi-Turn Text-to-SQL with Chain-of-EditionsabstractHanchong Zhang, Ruisheng Cao, Hongshen Xu, Lu Chen, Kai Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Hanchong Zhang, Ruisheng Cao, Hongshen Xu, Lu Chen 0002, Kai Yu 0004 |
NAACL-HLT | 4 |
| 2024 | Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?abstractData science and engineering workflows often span multiple stages, from warehousing to orchestration, using tools like BigQuery, dbt, and Airbyte. As vision language models (VLMs) advance in multimodal understanding and code generation, VLM-based agents could potentially automate these workflows by generating SQL queries, Python code, and GUI operations. This automation can improve the productivity of experts while democratizing access to large-scale data analysis. In this paper, we introduce Spider2-V, the first multimodal agent benchmark focusing on professional data science and engineering workflows, featuring 494 real-world tasks in authentic computer environments and incorporating 20 enterprise-level professional applications. These tasks, derived from real-world use cases, evaluate the ability of a multimodal agent to perform data-related tasks by writing code and managing the GUI in enterprise data software systems. To balance realistic simulation with evaluation simplicity, we devote significant effort to developing automatic configurations for task setup and carefully crafting evaluation metrics for each task. Furthermore, we supplement multimodal agents with comprehensive documents of these enterprise data software systems. Our empirical evaluation reveals that existing state-of-the-art LLM/VLM-based agents do not reliably automate full data workflows (14.0% success). Even with step-by-step guidance, these agents still underperform in tasks that require fine-grained, knowledge-intensive GUI actions (16.2%) and involve remote cloud-hosted workspaces (10.6%). We hope that Spider2-V paves the way for autonomous multimodal agents to transform the automation of data science and engineering workflow. Our code and data are available at https://spider2-v.github.io. Ruisheng Cao, Fangyu Lei, Haoyuan Wu, Jixuan Chen, Yeqiao Fu, Hongcheng Gao, Xinzhuang Xiong, Hanchong Zhang, Wenjing Hu, Tianbao Xie, Hongshen Xu, Sida I. Wang, Ruoxi Sun 0002, Caiming Xiong, Ansong Ni, Qian Liu 0033, Victor Zhong, Lu Chen 0002, Kai Yu 0004, Tao Yu 0009 |
NeurIPS | 21 |
| 2024 | Hierarchical Multimodal Pre-training for Visually Rich Webpage UnderstandingabstractThe growing prevalence of visually rich documents, such as webpages and scanned/digital-born documents (images, PDFs, etc.), has led to increased interest in automatic document understanding and information extraction across academia and industry. Although various document modalities, including image, text, layout, and structure, facilitate human information retrieval, the interconnected nature of these modalities presents challenges for neural networks. In this paper, we introduce WebLM, a multimodal pre-training network designed to address the limitations of solely modeling text and structure modalities of HTML in webpages. Instead of processing document images as unified natural images, WebLM integrates the hierarchical structure of document images to enhance the understanding of markup-language-based documents. Additionally, we propose several pre-training tasks to model the interaction among text, structure, and image modalities effectively. Empirical results demonstrate that the pre-trained WebLM significantly surpasses previous state-of-the-art pre-trained models across several webpage understanding tasks. The pre-trained models and code are available at https://github.com/X-LANCE/weblm. Hongshen Xu, Lu Chen 0002, Zihan Zhao 0001, Ruisheng Cao, Kai Yu 0004 |
WSDM | 2 |
| 2024 | ChemDFM-X: towards large multimodal model for chemistry
Zihan Zhao 0001, Jingpiao Li, Lu Chen 0002, Liyang Wen, Yansi Li, Zhongyang Dai, Kai Yu 0004 |
Sci. China Inf. Sci. | 4 |
| 2023 | How ChatGPT is Robust for Spoken Language Understanding?
Guangpeng Li, Lu Chen 0002, Kai Yu 0004 |
INTERSPEECH | 2 |
| 2023 | Large Language Models Are Semi-Parametric Reinforcement Learning AgentsabstractInspired by the insights in cognitive science with respect to human memory and reasoning mechanism, a novel evolvable LLM-based (Large Language Model) agent framework is proposed as Rememberer. By equipping the LLM with a long-term experience memory, Rememberer is capable of exploiting the experiences from the past episodes even for different task goals, which excels an LLM-based agent with fixed exemplars or equipped with a transient working memory. We further introduce **R**einforcement **L**earning with **E**xperience **M**emory (**RLEM**) to update the memory. Thus, the whole system can learn from the experiences of both success and failure, and evolve its capability without fine-tuning the parameters of the LLM. In this way, the proposed Rememberer constitutes a semi-parametric RL agent. Extensive experiments are conducted on two RL task sets to evaluate the proposed framework. The average results with different initialization and training sets exceed the prior SOTA by 4% and 2% for the success rate on two task sets and demonstrate the superiority and robustness of Rememberer. Lu Chen 0002, Situo Zhang, Hongshen Xu, Zihan Zhao 0001, Kai Yu 0004 |
NeurIPS | 2 |
| 2023 | A Heterogeneous Graph to Abstract Syntax Tree Framework for Text-to-SQLabstractText-to-SQL is the task of converting a natural language utterance plus the corresponding database schema into a SQL program. The inputs naturally form a heterogeneous graph while the output SQL can be transduced into an abstract syntax tree (AST). Traditional encoder-decoder models ignore higher-order semantics in heterogeneous graph encoding and introduce permutation biases during AST construction, thus incapable of exploiting the refined structure knowledge precisely. In this work, we propose a generic heterogeneous graph to abstract syntax tree (HG2AST) framework to integrate dedicated structure knowledge into statistics-based models. On the encoder side, we leverage a line graph enhanced encoder (LGESQL) to iteratively update both node and edge features through dual graph message passing and aggregation. On the decoder side, a grammar-based decoder first constructs the equivalent SQL AST and then transforms it into the desired SQL via post-processing. To avoid over-fitting permutation biases, we propose a golden tree-oriented learning (GTL) algorithm to adaptively control the expanding order of AST nodes. The graph encoder and tree decoder are combined into a unified framework through two auxiliary modules. Extensive experiments on various text-to-SQL datasets, including single/multi-table, single/cross-domain, and multilingual settings, demonstrate the superiority and broad applicability. Ruisheng Cao, Lu Chen 0002, Hanchong Zhang, Hongshen Xu, Wangyou Zhang, Kai Yu 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | OPAL: Ontology-Aware Pretrained Language Model for End-to-End Task-Oriented DialogueabstractAbstract This paper presents an ontology-aware pretrained language model (OPAL) for end-to-end task-oriented dialogue (TOD). Unlike chit-chat dialogue models, task-oriented dialogue models fulfill at least two task-specific modules: Dialogue state tracker (DST) and response generator (RG). The dialogue state consists of the domain-slot-value triples, which are regarded as the user’s constraints to search the domain-related databases. The large-scale task-oriented dialogue data with the annotated structured dialogue state usually are inaccessible. It prevents the development of the pretrained language model for the task-oriented dialogue. We propose a simple yet effective pretraining method to alleviate this problem, which consists of two pretraining phases. The first phase is to pretrain on large-scale contextual text data, where the structured information of the text is extracted by the information extracting tool. To bridge the gap between the pretraining method and downstream tasks, we design two pretraining tasks: ontology-like triple recovery and next-text generation, which simulates the DST and RG, respectively. The second phase is to fine-tune the pretrained model on the TOD data. The experimental results show that our proposed method achieves an exciting boost and obtains competitive performance even without any TOD data on CamRest676 and MultiWOZ benchmarks. Zhi Chen 0006, Yuncong Liu, Lu Chen 0002, Su Zhu, Mengyue Wu, Kai Yu 0004 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | AdapterShare: Task Correlation Modeling with Adapter DifferentiationabstractThanks to the development of pre-trained language models, multitask learning (MTL) methods have achieved great success in natural language understanding.However, current MTL methods pay more attention to task selection or model design to fuse as much knowledge as possible, while the intrinsic task correlation is often neglected.It is important to learn sharing strategies among multiple tasks rather than sharing everything.In this paper, we propose AdapterShare, an adapter differentiation method to explicitly model task correlation among multiple tasks.AdapterShare is automatically learned based on the gradients on tiny held-out validation data.Compared to single-task learning and fully shared MTL methods, our proposed method obtains obvious performance improvements.Compared to the existing MTL method AdapterFusion, AdapterShare achieves an absolute average improvement of 1.90 points on five dialogue understanding tasks and 2.33 points on NLU tasks.Our implementation is available at https:// github.com/microsoft/ContextualSP. Zhi Chen 0006, Bei Chen 0008, Lu Chen 0002, Kai Yu 0004, Jian-Guang Lou |
EMNLP | 3 |
| 2022 | META-GUI: Towards Multi-modal Conversational Agents on Mobile GUIabstractTask-oriented dialogue (TOD) systems have been widely used by mobile phone intelligent assistants to accomplish tasks such as calendar scheduling or hotel reservation.Current TOD systems usually focus on multi-turn text/speech interaction, then they would call back-end APIs designed for TODs to perform the task.However, this API-based architecture greatly limits the information-searching capability of intelligent assistants and may even lead to task failure if TOD-specific APIs are not available or the task is too complicated to be executed by the provided APIs.In this paper, we propose a new TOD architecture: GUI-based task-oriented dialogue system (GUI-TOD).A GUI-TOD system can directly perform GUI operations on real APPs and execute tasks without invoking TOD-specific backend APIs.Furthermore, we release META-GUI, a dataset for training a Multi-modal convErsaTional Agent on mobile GUI.We also propose a multi-model action prediction and response model, which show promising results on META-GUI.The dataset, codes and leaderboard are publicly available ‡ . Liangtai Sun, Lu Chen 0002, Tianle Dai, Kai Yu 0004 |
EMNLP | 3 |
| 2022 | D4: a Chinese Dialogue Dataset for Depression-Diagnosis-Oriented ChatabstractIn a depression-diagnosis-directed clinical session, doctors initiate a conversation with ample emotional support that guides the patients to expose their symptoms based on clinical diagnosis criteria.Such a dialogue system is distinguished from existing single-purpose humanmachine dialog systems, as it combines taskoriented and chit-chats with uniqueness in dialogue topics and procedures.However, due to the social stigma associated with mental illness, the dialogue data related to depression consultation and diagnosis are rarely disclosed.Based on clinical depression diagnostic criteria ICD-11 and DSM-5, we designed a 3phase procedure to construct D 4 : a Chinese Dialogue Dataset for Depression-Diagnosis-Oriented Chat 1 , which simulates the dialogue between doctors and patients during the diagnosis of depression, including diagnosis results and symptom summary given by professional psychiatrists for each conversation.Upon the newly-constructed dataset, four tasks mirroring the depression diagnosis process are established: response generation, topic prediction, dialog summary, and severity classification of depressive episode and suicide risk.Multiscale evaluation results demonstrate that a more empathy-driven and diagnostic-accurate consultation dialogue system trained on our dataset can be achieved compared to rule-based bots. Binwei Yao, Likai Zou, Lingfeng Dai, Mengyue Wu, Lu Chen 0002, Kai Yu 0004 |
EMNLP | 6 |
| 2022 | LatticeBART: Lattice-to-Lattice Pre-Training for Speech RecognitionabstractTo improve automatic speech recognition, increasing work has attempted to further fix the output of ASR systems with advanced sequence models. However, the output of ASR systems differs significantly from the input form of standard sequence models. To encompass richer information, the output of ASR systems is often a compact lattice structure containing multiple sentences. This mismatch in input form significantly limits sequence models’ ability. On the one hand, the widely used pre-trained models cannot directly input lattice structures and are therefore difficult to use for this task. On the other hand, the sparsity of the supervised training data forces the model to have the ability to learn from limited data. To address these problems, we propose LatticeBART, a model that decodes the sequence from the lattice in an end-to-end fashion and can use the pre-trained language models’ prior. In addition, this paper proposes the lattice-to-lattice pre-training method, which can be used when annotated data is missing, using easily generated lattice with the ASR system for training. The experimental results show that our model can effectively improve the output quality of the ASR system. Lingfeng Dai, Lu Chen 0002, Zhikai Zhou, Kai Yu 0004 |
ICASSP | 2 |
| 2022 | TIE: Topological Information Enhanced Structural Reading Comprehension on Web PagesabstractZihan Zhao, Lu Chen, Ruisheng Cao, Hongshen Xu, Xingyu Chen, Kai Yu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Zihan Zhao 0001, Lu Chen 0002, Ruisheng Cao, Hongshen Xu, Kai Yu 0004 |
NAACL-HLT | 2 |
| 2022 | UniDU: Towards A Unified Generative Dialogue Understanding FrameworkabstractWith the development of pre-trained language models, remarkable success has been witnessed in dialogue understanding (DU).However, current DU approaches usually employ independent models for each distinct DU task without considering shared knowledge across different DU tasks.In this paper, we propose a unified generative dialogue understanding framework, named UniDU, to achieve effective information exchange across diverse DU tasks.Here, we reformulate all DU tasks into a unified promptbased generative model paradigm.More importantly, a novel model-agnostic multi-task training strategy (MATS) is introduced to dynamically adapt the weights of diverse tasks for best knowledge sharing during training, based on the nature and available data of each task.Experiments on ten DU datasets covering five fundamental DU tasks show that the proposed UniDU framework largely outperforms task-specific well-designed methods on all tasks.MATS also reveals the knowledgesharing structure of these tasks.Finally, UniDU obtains promising performance in the unseen dialogue domain, showing the great potential for generalization. Zhi Chen 0006, Lu Chen 0002, Bei Chen 0008, Libo Qin 0001, Yuncong Liu, Su Zhu, Jian-Guang Lou, Kai Yu 0004 |
SIGDIAL | 2 |
| 2021 | LET: Linguistic Knowledge Enhanced Graph Transformer for Chinese Short Text MatchingabstractChinese short text matching is a fundamental task in natural language processing. Existing approaches usually take Chinese characters or words as input tokens. They have two limitations: 1) Some Chinese words are polysemous, and semantic information is not fully utilized. 2) Some models suffer potential issues caused by word segmentation. Here we introduce HowNet as an external knowledge base and propose a Linguistic knowledge Enhanced graph Transformer (LET) to deal with word ambiguity. Additionally, we adopt the word lattice graph as input to maintain multi-granularity information. Our model is also complementary to pre-trained language models. Experimental results on two Chinese datasets show that our models outperform various typical text matching approaches. Ablation study also indicates that both semantic information and multi-granularity information are important for text matching modeling. Boer Lyu, Lu Chen 0002, Su Zhu, Kai Yu 0004 |
AAAI | 2 |
| 2021 | LGESQL: Line Graph Enhanced Text-to-SQL Model with Mixed Local and Non-Local RelationsabstractRuisheng Cao, Lu Chen, Zhi Chen, Yanbin Zhao, Su Zhu, Kai Yu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ruisheng Cao, Lu Chen 0002, Zhi Chen 0006, Yanbin Zhao, Su Zhu, Kai Yu 0004 |
ACL/IJCNLP (1) | 2 |
| 2021 | WebSRC: A Dataset for Web-Based Structural Reading ComprehensionabstractWeb search is an essential way for humans to obtain information, but it's still a great challenge for machines to understand the contents of web pages.In this paper, we introduce the task of structural reading comprehension (SRC) on web.Given a web page and a question about it, the task is to find the answer from the web page.This task requires a system not only to understand the semantics of texts but also the structure of the web page.Moreover, we proposed Web-SRC, a novel Web-based Structural Reading Comprehension dataset.WebSRC consists of 400K question-answer pairs, which are collected from 6.4K web pages.Along with the QA pairs, corresponding HTML source code, screenshots, and metadata are also provided in our dataset.Each question in WebSRC requires a certain structural understanding of a web page to answer, and the answer is either a text span on the web page or yes/no.We evaluate various baselines on our dataset to show the difficulty of our task.We also investigate the usefulness of structural information and visual features.Our dataset and baselines have been publicly available 1 . Zihan Zhao 0001, Lu Chen 0002, Jiabao Ji, Ao Luo, Yuxuan Xiong, Kai Yu 0004 |
EMNLP (1) | 3 |
| 2021 | ShadowGNN: Graph Projection Neural Network for Text-to-SQL ParserabstractZhi Chen, Lu Chen, Yanbin Zhao, Ruisheng Cao, Zihan Xu, Su Zhu, Kai Yu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhi Chen 0006, Lu Chen 0002, Yanbin Zhao, Ruisheng Cao, Su Zhu, Kai Yu 0004 |
NAACL-HLT | 2 |
| 2021 | Relation-Aware Multi-hop Reasoning forVisual Dialog
Lu Chen 0002, Kai Yu 0004 |
NLPCC (1) | 2 |
| 2021 | Few-Shot NLU with Vector Projection Distance and Abstract Triangular CRF
Su Zhu, Lu Chen 0002, Ruisheng Cao, Zhi Chen 0006, Qingliang Miao, Kai Yu 0004 |
NLPCC (1) | 2 |
| 2020 | Schema-Guided Multi-Domain Dialogue State Tracking with Graph Attention Neural NetworksabstractDialogue state tracking (DST) aims at estimating the current dialogue state given all the preceding conversation. For multi-domain DST, the data sparsity problem is also a major obstacle due to the increased number of state candidates. Existing approaches generally predict the value for each slot independently and do not consider slot relations, which may aggravate the data sparsity problem. In this paper, we propose a Schema-guided multi-domain dialogue State Tracker with graph attention networks (SST) that predicts dialogue states from dialogue utterances and schema graphs which contain slot relations in edges. We also introduce a graph attention matching network to fuse information from utterances and graphs, and a recurrent graph attention network to control state updating. Experiment results show that our approach obtains new state-of-the-art performance on both MultiWOZ 2.0 and MultiWOZ 2.1 benchmarks. Lu Chen 0002, Boer Lv, Su Zhu, Bowen Tan, Kai Yu 0004 |
AAAI | 1 |
| 2020 | Semi-Supervised Text Simplification with Back-Translation and Asymmetric Denoising AutoencodersabstractText simplification (TS) rephrases long sentences into simplified variants while preserving inherent semantics. Traditional sequence-to-sequence models heavily rely on the quantity and quality of parallel sentences, which limits their applicability in different languages and domains. This work investigates how to leverage large amounts of unpaired corpora in TS task. We adopt the back-translation architecture in unsupervised machine translation (NMT), including denoising autoencoders for language modeling and automatic generation of parallel data by iterative back-translation. However, it is non-trivial to generate appropriate complex-simple pair if we directly treat the set of simple and complex corpora as two different languages, since the two types of sentences are quite similar and it is hard for the model to capture the characteristics in different types of sentences. To tackle this problem, we propose asymmetric denoising methods for sentences with separate complexity. When modeling simple and complex sentences with autoencoders, we introduce different types of noise into the training process. Such a method can significantly improve the simplification performance. Our model can be trained in both unsupervised and semi-supervised manner. Automatic and human evaluations show that our unsupervised model outperforms the previous systems, and with limited supervision, our model can perform competitively with multiple state-of-the-art simplification systems. Yanbin Zhao, Lu Chen 0002, Zhi Chen 0006, Kai Yu 0004 |
AAAI | 2 |
| 2020 | Unsupervised Dual Paraphrasing for Two-stage Semantic ParsingabstractOne daunting problem for semantic parsing is the scarcity of annotation.Aiming to reduce nontrivial human labor, we propose a two-stage semantic parsing framework, where the first stage utilizes an unsupervised paraphrase model to convert an unlabeled natural language utterance into the canonical utterance.The downstream naive semantic parser accepts the intermediate output and returns the target logical form.Furthermore, the entire training process is split into two phases: pre-training and cycle learning.Three tailored self-supervised tasks are introduced throughout training to activate the unsupervised paraphrase model.Experimental results on benchmarks OVERNIGHT and GE-OGRANNO demonstrate that our framework is effective and compatible with supervised training. Ruisheng Cao, Su Zhu, Chen Liu 0019, Rao Ma, Yanbin Zhao, Lu Chen 0002, Kai Yu 0004 |
ACL | 7 |
| 2020 | Neural Graph Matching Networks for Chinese Short Text MatchingabstractChinese short text matching usually employs word sequences rather than character sequences to get better performance.However, Chinese word segmentation can be erroneous, ambiguous or inconsistent, which consequently hurts the final matching performance.To address this problem, we propose neural graph matching networks, a novel sentence matching framework capable of dealing with multi-granular input information.Instead of a character sequence or a single word sequence, paired word lattices formed from multiple word segmentation hypotheses are used as input and the model learns a graph representation according to an attentive graph matching mechanism.Experiments on two Chinese datasets show that our models outperform the state-of-the-art short text matching models. Lu Chen 0002, Yanbin Zhao, Boer Lyu, Lesheng Jin, Zhi Chen 0006, Su Zhu, Kai Yu 0004 |
ACL | 1 |
| 2020 | Line Graph Enhanced AMR-to-Text Generation with Mix-Order Graph Attention NetworksabstractEfficient structure encoding for graphs with labeled edges is an important yet challenging point in many graph-based models.This work focuses on AMR-to-text generation -A graph-to-sequence task aiming to recover natural language from Abstract Meaning Representations (AMR).Existing graph-to-sequence approaches generally utilize graph neural networks as their encoders, which have two limitations: 1) The message propagation process in AMR graphs is only guided by the firstorder adjacency information.2) The relationships between labeled edges are not fully considered.In this work, we propose a novel graph encoding framework which can effectively explore the edge relations.We also adopt graph attention networks with higherorder neighborhood information to encode the rich structure in AMR graphs.Experiment results show that our approach obtains new state-of-the-art performance on English AMR benchmark datasets.The ablation analyses also demonstrate that both edge relations and higher-order information are beneficial to graph-to-sequence modeling. Yanbin Zhao, Lu Chen 0002, Zhi Chen 0006, Ruisheng Cao, Su Zhu, Kai Yu 0004 |
ACL | 2 |
| 2020 | Addressing the Polysemy Problem in Language Modeling with Attentional Multi-Sense EmbeddingsabstractNeural network language models have gained considerable popularity due to their promising performance. Distributed word embeddings are utilized to represent semantic information. However, each word is associated with a single vector in the embedding layer, disabling the model from capturing the meanings of polysemous words. In this work, we address this problem by assigning multiple fine-grained sense embeddings to each word in the embedding layers. The proposed model discriminates among different senses of a word with attention mechanism in an unsupervised manner. Experiments demonstrate the benefits of our approach in language modeling and ASR rescoring. Investigations are also made on standard word similarity tasks. The results indicate that our proposed method is efficient in modeling polysemy and therefore obtains better word representations. Rao Ma, Lesheng Jin, Qi Liu 0018, Lu Chen 0002, Kai Yu 0004 |
ICASSP | 4 |
| 2020 | Neural Lattice Search for Speech RecognitionabstractTo improve the accuracy of automatic speech recognition, a two-pass decoding strategy is widely adopted. The first-pass model generates compact word lattices, which are utilized by the second-pass model to perform rescoring. Currently, the most popular rescoring methods are N-best rescoring and lattice rescoring with long short-term memory language models (LSTMLMs). However, these methods encounter the problem of limited search space or inconsistency between training and evaluation. In this paper, we address these problems with an end-to-end model for accurately extracting the best hypothesis from the word lattice. Our model is composed of a bidirectional LatticeLSTM encoder followed by an attentional LSTM decoder. The model takes word lattice as input and generates the single best hypothesis from the given lattice space. When combined with an LSTMLM, the proposed model yields 9.7% and 7.5% relative WER reduction compared to N-best rescoring methods and lattice rescoring methods within the same amount of decoding time. Rao Ma, Qi Liu 0018, Lu Chen 0002, Kai Yu 0004 |
ICASSP | 4 |
| 2020 | Jointly Encoding Word Confusion Network and Dialogue Context with BERT for Spoken Language UnderstandingabstractSpoken Language Understanding (SLU) converts hypotheses from automatic speech recognizer (ASR) into structured semantic representations. ASR recognition errors can severely degenerate the performance of the subsequent SLU module. To address this issue, word confusion networks (WCNs) have been used to encode the input for SLU, which contain richer information than 1-best or n-best hypotheses list. To further eliminate ambiguity, the last system act of dialogue context is also utilized as additional input. In this paper, a novel BERT based SLU model (WCN-BERT SLU) is proposed to encode WCNs and the dialogue context jointly. It can integrate both structural information and ASR posterior probabilities of WCNs in the BERT architecture. Experiments on DSTC2, a benchmark of SLU, show that the proposed method is effective and can outperform previous state-of-the-art models significantly. Chen Liu 0019, Su Zhu, Ruisheng Cao, Lu Chen 0002, Kai Yu 0004 |
INTERSPEECH | 5 |
| 2020 | Robust Spoken Language Understanding with RL-Based Value Error Recovery
Chen Liu 0019, Su Zhu, Lu Chen 0002, Kai Yu 0004 |
NLPCC (1) | 3 |
| 2020 | Memory Attention Neural Network for Multi-domain Dialogue State Tracking
Zhi Chen 0006, Lu Chen 0002, Su Zhu, Kai Yu 0004 |
NLPCC (1) | 3 |
| 2020 | An Investigation on Different Underlying Quantization Schemes for Pre-trained Language Models
Zihan Zhao 0001, Yuncong Liu, Lu Chen 0002, Qi Liu 0018, Rao Ma, Kai Yu 0004 |
NLPCC (1) | 3 |
| 2020 | Distributed Structured Actor-Critic Reinforcement Learning for Universal Dialogue ManagementabstractTraditional dialogue policy needs to be trained independently for each dialogue task. In this work, we aim to solve a collection of independent dialogue tasks using a unified dialogue agent. The unified policy is parallelly trained using the conversation data from all of the distributed dialogue tasks. However, there are two key challenges:(1) the design of a unified dialogue model to adapt to different dialogue tasks; (2) finding a robust reinforcement learning method to keep the efficiency and the stability of the training process. Here we propose a novel structured actor-critic approach to implement structured deep reinforcement learning (DRL), which not only can learn parallelly from data of different dialogue tasks but also achieves stable and sample-efficient learning. We demonstrate the effectiveness of the proposed approach on 18 tasks of PyDial benchmark. The results show that our method is able to achieve state-of-the-art performance. Zhi Chen 0006, Lu Chen 0002, Kai Yu 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | AgentGraph: Toward Universal Dialogue Management With Structured Deep Reinforcement LearningabstractDialogue policy plays an important role in task-oriented spoken dialogue systems. It determines how to respond to users. The recently proposed deep reinforcement learning (DRL) approaches have been used for policy optimization. However, these deep models are still challenging for two reasons: first, many DRL-based policies are not sample efficient; and second, most models do not have the capability of policy transfer between different domains. In this paper, we propose a universal framework, AgentGraph, to tackle these two problems. The proposed AgentGraph is the combination of graph neural network (GNN) based architecture and DRL-based algorithm. It can be regarded as one of the multi-agent reinforcement learning approaches. Each agent corresponds to a node in a graph, which is defined according to the dialogue domain ontology. When making a decision, each agent can communicate with its neighbors on the graph. Under AgentGraph framework, we further propose dual GNN-based dialogue policy, which implicitly decomposes the decision in each turn into a high-level global decision and a low-level local decision. Experiments show that AgentGraph models significantly outperform traditional reinforcement learning approaches on most of the 18 tasks of the PyDial benchmark. Moreover, when transferred from the source task to a target task, these models not only have acceptable initial performance but also converge much faster on the target task. Lu Chen 0002, Zhi Chen 0006, Bowen Tan, Sishan Long, Milica Gasic, Kai Yu 0004 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Structured Dialogue Policy with Graph Neural NetworksabstractRecently, deep reinforcement learning (DRL) has been used for dialogue policy optimization. However, many DRL-based policies are not sample-efficient. Most recent advances focus on improving DRL optimization algorithms to address this issue. Here, we take an alternative route of designing neural network structure that is better suited for DRL-based dialogue management. The proposed structured deep reinforcement learning is based on graph neural networks (GNN), which consists of some sub-networks, each one for a node on a directed graph. The graph is defined according to the domain ontology and each node can be considered as a sub-agent. During decision making, these sub-agents have internal message exchange between neighbors on the graph. We also propose an approach to jointly optimize the graph structure as well as the parameters of GNN. Experiments show that structured DRL significantly outperforms previous state-of-the-art approaches in almost all of the 18 tasks of the PyDial benchmark. Lu Chen 0002, Bowen Tan, Sishan Long, Kai Yu 0004 |
COLING | 1 |
| 2018 | Towards Universal Dialogue State TrackingabstractDialogue state tracking is the core part of a spoken dialogue system.It estimates the beliefs of possible user's goals at every dialogue turn.However, for most current approaches, it's difficult to scale to large dialogue domains.They have one or more of following limitations: ( a) Some models don't work in the situation where slot values in ontology changes dynamically; (b) The number of model parameters is proportional to the number of slots; (c) Some models extract features based on hand-crafted lexicons.To tackle these challenges, we propose StateNet, a universal dialogue state tracker.It is independent of the number of values, shares parameters across all slots, and uses pre-trained word vectors instead of explicit semantic dictionaries.Our experiments on two datasets show that our approach not only overcomes the limitations, but also significantly outperforms the performance of state-of-the-art approaches. Liliang Ren, Kaige Xie, Lu Chen 0002, Kai Yu 0004 |
EMNLP | 3 |
| 2018 | Policy Adaptation for Deep Reinforcement Learning-Based Dialogue ManagementabstractPolicy optimization is the core part of statistical dialogue management. Deep reinforcement learning has been successfully used for dialogue policy optimization for a static pre-defined domain. However, when the domain changes dynamically, e.g. a new previously unseen concept (or slot) which can be then used as a database search constraint is added, or the policy for one domain is transferred to another domain, the dialogue state space and action sets both will change. Therefore, the model structures for different domains have to be different. This makes dialogue policy adaptation/transfer challenging. Here a multi -agent dialogue policy (MADP) is proposed to tackle these problems. MADP consists of some slot-dependent agents (S-Agents) and a slot-independent agent (G-Agent). S-Agents have shared parameters in addition to private parameters for each one. During policy transfer, the shared parameters in S-Agents and the parameters in G-Agent can be directly transferred to the agents in extended/new domain. Simulation experiments showed that MADP can significantly speed up the policy learning and facilitate policy adaptation. Lu Chen 0002, Zhi Chen 0006, Bowen Tan, Milica Gasic, Kai Yu 0004 |
ICASSP | 1 |
| 2018 | Cost-Sensitive Active Learning for Dialogue State TrackingabstractDialogue state tracking (DST), when formulated as a supervised learning problem, relies on labelled data.Since dialogue state annotation usually requires labelling all turns of a single dialogue and utilizing context information, it is very expensive to annotate all available unlabelled data.In this paper, a novel cost-sensitive active learning framework is proposed based on a set of new dialogue-level query strategies.This is the first attempt to apply active learning for dialogue state tracking.Experiments on DSTC2 show that active learning with mixed data query strategies can effectively achieve the same DST performance with significantly less data annotation compared to traditional training approaches. Kaige Xie, Liliang Ren, Lu Chen 0002, Kai Yu 0004 |
SIGDIAL Conference | 4 |
| 2017 | Affordable On-line Dialogue Policy LearningabstractThe key to building an evolvable dialogue system in real-world scenarios is to ensure an affordable on-line dialogue policy learning, which requires the on-line learning process to be safe, efficient and economical.But in reality, due to the scarcity of real interaction data, the dialogue system usually grows slowly.Besides, the poor initial dialogue policy easily leads to bad user experience and incurs a failure of attracting users to contribute training data, so that the learning process is unsustainable.To accurately depict this, two quantitative metrics are proposed to assess safety and efficiency issues.For solving the unsustainable learning problem, we proposed a complete companion teaching framework incorporating the guidance from the human teacher.Since the human teaching is expensive, we compared various teaching schemes answering the question how and when to teach, to economically utilize teaching budget, so that make the online learning process affordable. Runzhe Yang, Lu Chen 0002, Kai Yu 0004 |
EMNLP | 3 |
| 2017 | Agent-Aware Dropout DQN for Safe and Efficient On-line Dialogue Policy LearningabstractHand-crafted rules and reinforcement learning (RL) are two popular choices to obtain dialogue policy.The rule-based policy is often reliable within predefined scope but not self-adaptable, whereas RL is evolvable with data but often suffers from a bad initial performance.We employ a companion learning framework to integrate the two approaches for on-line dialogue policy learning, in which a predefined rule-based policy acts as a teacher and guides a data-driven RL system by giving example actions as well as additional rewards.A novel agent-aware dropout Deep Q-Network (AAD-DQN) is proposed to address the problem of when to consult the teacher and how to learn from the teacher's experiences.AAD-DQN, as a data-driven student policy, provides (1) two separate experience memories for student and teacher, (2) an uncertainty estimated by dropout to control the timing of consultation and learning.Simulation experiments showed that the proposed approach can significantly improve both safety and efficiency of on-line policy optimization compared to other companion learning approaches as well as supervised pre-training using static dialogue corpus. Lu Chen 0002, Runzhe Yang, Kai Yu 0004 |
EMNLP | 1 |
| 2016 | Hybrid Dialogue State Tracking for Real World Human-to-Human Dialogues
Su Zhu, Lu Chen 0002, Siqiu Yao, Xueyang Wu 0001, Kai Yu 0004 |
INTERSPEECH | 3 |
| 2016 | Evolvable dialogue state tracking for statistical dialogue management
Kai Yu 0004, Lu Chen 0002, Qizhe Xie, Su Zhu |
Frontiers Comput. Sci. | 2 |
| 2015 | Hyper-parameter Optimisation of Gaussian Process Reinforcement Learning for Statistical Dialogue ManagementabstractGaussian processes reinforcement learning provides an appealing framework for training the dialogue policy as it takes into account correlations of the objective function given different dialogue belief states, which can significantly speed up the learning.These correlations are modelled by the kernel function which may depend on hyper-parameters.So far, for real-world dialogue systems the hyperparameters have been hand-tuned, relying on the designer to adjust the correlations, or simple non-parametrised kernel functions have been used instead.Here, we examine different kernel structures and show that it is possible to optimise the hyperparameters from data yielding improved performance of the resulting dialogue policy.We confirm this in a real user trial. Lu Chen 0002, Pei-hao Su, Milica Gasic |
SIGDIAL Conference | 1 |
| 2015 | Recurrent Polynomial Network for Dialogue State Tracking with Mismatched Semantic ParsersabstractRecently, constrained Markov Bayesian polynomial (CMBP) has been proposed as a data-driven rule-based model for dialog state tracking (DST).CMBP is an approach to bridge rule-based models and statistical models.Recurrent Polynomial Network (RPN) is a recent statistical framework taking advantages of rulebased models and can achieve state-ofthe-art performance on the data corpora of DSTC-3, outperforming all submitted trackers in DSTC-3 including RNN.It is widely acknowledged that SLU's reliability influences tracker's performance greatly, especially in cases where the training SLU is poorly matched to the testing SLU.In this paper, this effect is analyzed in detail for RPN.Experiments show that RPN's tracking result is consistently the best compared to rule-based and statistical models investigated on different SLUs including mismatched ones and demonstrate RPN's is very robust to mismatched semantic parsers. Qizhe Xie, Su Zhu, Lu Chen 0002, Kai Yu 0004 |
SIGDIAL Conference | 4 |
| 2015 | Constrained Markov Bayesian Polynomial for Efficient Dialogue State TrackingabstractDialogue state tracking (DST) is a process to estimate the distribution of the dialogue states at each dialogue turn given the interaction history. Although data-driven statistical approaches are of most interest, there have been attempts of using rule-based methods for DST, due to their simplicity, efficiency and portability. However, the performance of these methods are usually not competitive to data-driven tracking approaches and it is not possible to improve the DST performance when training data are available. In this paper, a novel hybrid framework, constrained Markov Bayesian polynomial (CMBP), is proposed to formulate rule-based DST in a general way and allow data-driven rule generation. Here, a DST rule is defined as a polynomial function of a set of probabilities satisfying certain linear constraints. Prior knowledge is encoded in these constraints. Under reasonable assumptions, CMBP optimization can be converted to a constrained integer linear programming problem. The integer coefficient CMBP model is further extended to CMBP with real coefficients by applying grid search. CMBP was evaluated on the data corpora of the first, the second, and the third Dialog State Tracking Challenge (DSTC-1/2/3). Experiments showed that CMBP has good generalization ability and can significantly outperform both traditional rule-based approaches and data-driven statistical approaches with similar feature set. Compared with the state-of-the-art statistical DST approaches with much richer features, CMBP is also competitive. Kai Yu 0004, Lu Chen 0002, Su Zhu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | The SJTU System for Dialog State Tracking Challenge 2abstractDialog state tracking challenge provides a common testbed for state tracking al-gorithms. This paper describes the SJTU system submitted to the second Dialogue State Tracking Challenge in detail. In the system, a statistical semantic parser is used to generate refined semantic hypothe-ses. A large number of features are then derived based on the semantic hypothe-ses and the dialogue log information. The final tracker is a combination of a rule-based model, a maximum entropy and a deep neural network model. The SJTU system significantly outperformed all the baselines and showed competitive perfor-mance in DSTC 2. 1 Lu Chen 0002, Su Zhu, Kai Yu 0004 |
SIGDIAL Conference | 2 |
| 2014 | A generalized rule based tracker for dialogue state trackingabstractDialogue state tracking plays an important role in statistical dialogue management. Domain-independent rule-based approaches are attractive due to their efficiency, portability and interpretability. However, recent rule-based models are still not quite competitive to statistical tracking approaches. In this paper, a novel framework is proposed to formulate rule-based models in a general way. In the framework, a rule is considered as a special kind of polynomial function satisfying certain linear constraints. Under some particular definitions and assumptions, rule-based models can be seen as feasible solutions of an integer linear programming problem. Experiments showed that the proposed approach can not only achieve competitive performance compared to statistical approaches, but also have good generalisation ability. It is one of the only two entries that outperformed all the four baselines in the third Dialog State Tracking Challenge. Lu Chen 0002, Su Zhu, Kai Yu 0004 |
SLT | 2 |
| 2014 | Semantic parser enhancement for dialogue domain extension with little dataabstractStatistical semantic parser trained on sufficient in-domain data has shown robustness to speech recognition errors in end-to-end spoken dialogue systems. However, when the dialogue domain is extended, due to the introduction of new semantic slots, values and unknown speech pattern, the parsing performance may significantly degrade. Effective re-training of statistical semantic parser is therefore important. This paper describes a novel semantic parser enhancement approach for domain extension with very little new data. It employs automatic pseudo-data generation for parser re-training and domain independent rescoring to further improve parsing performance. The approach was evaluated on the DSTC3 (the third Dialog State Tracking Challenge) data corpus. Experiments showed that the proposed approach can yield consistent and significant improvements across all metrics of semantic parsing and dialog state tracking. Su Zhu, Lu Chen 0002, Kai Yu 0004 |
SLT | 2 |