VLDB 2026 Research / reviewers in the wild / expert
Jia-Chen Gu
dblp:93/3604
· DBLP profile ↗
30ranked-venue papers
14as first author
23since 2021 · last 2026
0000-0002-8801-1438ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 12 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can Editing LLMs Inject Harm?abstractLarge Language Models (LLMs) have emerged as a new information channel. Meanwhile, one critical but under-explored question is: Is it possible to bypass the safety alignment and inject harmful information into LLMs stealthily? In this paper, we propose to reformulate knowledge editing as a new type of safety threat for LLMs, namely Editing Attack, and conduct a systematic investigation with a newly constructed dataset EditAttack. Specifically, we focus on two typical safety risks of Editing Attack including Misinformation Injection and Bias Injection. For the first risk, we find that editing attacks can inject both commonsense and long-tail misinformation into LLMs, and the effectiveness for the former one is particularly high. For the second risk, we discover that not only can biased sentences be injected into LLMs with high effectiveness, but also one single biased sentence injection can degrade the overall fairness. Then, we further illustrate the high stealthiness of editing attacks. Our discoveries demonstrate the emerging misuse risks of knowledge editing techniques on compromising the safety alignment of LLMs and the feasibility of disseminating misinformation or bias with LLMs as new channels. Canyu Chen, Baixiang Huang, Zekun Li 0001, Zhaorun Chen, Shiyang Lai, Xiongxiao Xu, Jia-Chen Gu, Jindong Gu, Huaxiu Yao, Chaowei Xiao, Xifeng Yan, William Yang Wang, Philip Torr 0001, Dawn Song, Kai Shu |
AAAI | 7 |
| 2026 | Multiplicative Orthogonal Sequential Editing for Language ModelsabstractKnowledge editing aims to efficiently modify the internal knowledge of large language models (LLMs) without compromising their other capabilities. The prevailing editing paradigm, which appends an update matrix to the original parameter matrix, has been shown by some studies to damage key numerical stability indicators (such as condition number and norm), thereby reducing editing performance and general abilities, especially in sequential editing scenario. Although subsequent methods have made some improvements, they remain within the additive framework and have not fundamentally addressed this limitation. To solve this problem, we analyze it from both statistical and mathematical perspectives and conclude that multiplying the original matrix by an orthogonal matrix does not change the numerical stability of the matrix. Inspired by this, different from the previous additive editing paradigm, a multiplicative editing paradigm termed Multiplicative Orthogonal Sequential Editing (MOSE) is proposed. Specifically, we first derive the matrix update in the multiplicative form, the new knowledge is then incorporated into an orthogonal matrix, which is multiplied by the original parameter matrix. In this way, the numerical stability of the edited matrix is unchanged, thereby maintaining editing performance and general abilities. We compared MOSE with several current knowledge editing methods, systematically evaluating their impact on both editing performance and the general abilities across three different LLMs. Experimental results show that MOSE effectively limits deviations in the edited parameter matrix and maintains its numerical stability. Compared to current methods, MOSE achieves a 12.08% improvement in sequential editing performance, while retaining 95.73% of general abilities across downstream tasks. Hao-Xiang Xu, Jun-Yu Ma, Ziqi Peng, Zhen-Hua Ling, Jia-Chen Gu |
AAAI | 6 |
| 2025 | MISP-Meeting: A Real-World Dataset with Multimodal Cues for Long-form Meeting Transcription and SummarizationabstractHangChen HangChen, Chao-Han Huck Yang, Jia-Chen Gu, Sabato Marco Siniscalchi, Jun Du. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. HangChen HangChen, Chao-Han Huck Yang, Jia-Chen Gu, Sabato Marco Siniscalchi, Jun Du 0002 |
ACL (1) | 3 |
| 2025 | CaKE: Circuit-aware Editing Enables Generalizable Knowledge LearnersabstractKnowledge Editing (KE) enables the modification of outdated or incorrect information in large language models (LLMs).While existing KE methods can update isolated facts, they often fail to generalize these updates to multihop reasoning tasks that rely on the modified knowledge.Through an analysis of reasoning circuits-the neural pathways LLMs use for knowledge-based inference, we find that current layer-localized KE approaches (e.g., MEMIT, WISE), which edit only single or a few model layers, inadequately integrate updated knowledge into these reasoning pathways.To address this limitation, we present CaKE (Circuit-aware Knowledge Editing), a novel method that enhances the effective integration of updated knowledge in LLMs.By only leveraging a few curated data samples guided by our circuit-based analysis, CaKE stimulates the model to develop appropriate reasoning circuits for newly incorporated knowledge.Experiments show that CaKE enables more accurate and consistent use of edited knowledge across related reasoning tasks, achieving an average improvement of 20% in multi-hop reasoning accuracy on the MQuAKE dataset while requiring less memory than existing KE methods.We release the code and data in https://github.com/zjunlp/CaKE. Yunzhi Yao, Jizhan Fang, Jia-Chen Gu, Ningyu Zhang 0001, Shumin Deng, Huajun Chen, Nanyun Peng 0001 |
EMNLP | 3 |
| 2025 | MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal ModelsabstractExisting multimodal retrieval benchmarks primarily focus on evaluating whether models can retrieve and utilize external textual knowledge for question answering. However, there are scenarios where retrieving visual information is either more beneficial or easier to access than textual data.
In this paper, we introduce a multimodal retrieval-augmented generation benchmark, MRAG-Bench, in which we systematically identify and categorize scenarios where visually augmented knowledge is better than textual knowledge, for instance, more images from varying viewpoints.
MRAG-Bench consists of 16,130 images and 1,353 human-annotated multiple-choice questions across 9 distinct scenarios. With MRAG-Bench, we conduct an evaluation of 10 open-source and 4 proprietary large vision-language models (LVLMs). Our results show that all LVLMs exhibit greater improvements when augmented with images compared to textual knowledge, confirming that MRAG-Bench is vision-centric. Additionally, we conduct extensive analysis with MRAG-Bench, which offers valuable insights into retrieval-augmented LVLMs. Notably, the top-performing model, GPT-4o, faces challenges in effectively leveraging retrieved knowledge, achieving only a 5.82\% improvement with ground-truth information, in contrast to a 33.16\% improvement observed in human participants. These findings highlight the importance of MRAG-Bench in encouraging the community to enhance LVLMs' ability to utilize retrieved visual knowledge more effectively. Wenbo Hu 0006, Jia-Chen Gu, Zi-Yi Dou, Mohsen Fayyaz, Pan Lu, Kai-Wei Chang 0001, Nanyun Peng 0001 |
ICLR | 2 |
| 2025 | Perturbation-Restrained Sequential Model EditingabstractModel editing is an emerging field that focuses on updating the knowledge embedded within large language models (LLMs) without extensive retraining. However, current model editing methods significantly compromise the general abilities of LLMs as the number of edits increases, and this trade-off poses a substantial challenge to the continual learning of LLMs. In this paper, we first theoretically analyze that the factor affecting the general abilities in sequential model editing lies in the condition number of the edited matrix. The condition number of a matrix represents its numerical sensitivity, and therefore can be used to indicate the extent to which the original knowledge associations stored in LLMs are perturbed after editing. Subsequently, statistical findings demonstrate that the value of this factor becomes larger as the number of edits increases, thereby exacerbating the deterioration of general abilities. To this end, a framework termed Perturbation Restraint on Upper bouNd for Editing (PRUNE) is proposed, which applies the condition number restraints in sequential editing. These restraints can lower the upper bound on perturbation to edited models, thus preserving the general abilities.
Systematically, we conduct experiments employing three editing methods on three LLMs across four downstream tasks.
The results show that PRUNE can preserve general abilities while maintaining the editing performance effectively in sequential model editing. The code are available at https://github.com/mjy1111/PRUNE. Jun-Yu Ma, Hong Wang 0028, Hao-Xiang Xu, Zhen-Hua Ling, Jia-Chen Gu |
ICLR | 5 |
| 2025 | Accelerating Inference of Retrieval-Augmented Generation via Sparse Context SelectionabstractLarge language models (LLMs) augmented with retrieval exhibit robust performance and extensive versatility by incorporating external contexts. However, the input length grows linearly in the number of retrieved documents, causing a dramatic increase in latency. In this paper, we propose a novel paradigm named Sparse RAG, which seeks to cut computation costs through sparsity. Specifically, Sparse RAG encodes retrieved documents in parallel, which eliminates latency introduced by long-range attention of retrieved documents. Then, LLMs selectively decode the output by only attending to highly relevant caches auto-regressively, which are chosen via prompting LLMs with special control tokens. It is notable that Sparse RAG combines the assessment of each individual document and the generation of the response into a single process. The designed sparse mechanism in a RAG system can facilitate the reduction of the number of documents loaded during decoding for accelerating the inference of the RAG system. Additionally, filtering out undesirable contexts enhances the model’s focus on relevant context, inherently improving its generation quality. Evaluation results on four datasets show that Sparse RAG can be used to strike an optimal balance between generation quality and computational efficiency, demonstrating its generalizability across tasks. Jia-Chen Gu, Caitlin Sikora, Ho Ko, Yinxiao Liu, Chu-Cheng Lin, Lei Shu 0004, Liangchen Luo, Lei Meng 0008, Jindong Chen |
ICLR | 2 |
| 2024 | Model Editing Harms General Abilities of Large Language Models: Regularization to the RescueabstractModel editing is a technique that edits the large language models (LLMs) with updated knowledge to alleviate hallucinations without resource-intensive retraining.While current model editing methods can effectively modify a model's behavior within a specific area of interest, they often overlook the potential unintended side effects on the general abilities of LLMs such as reasoning, natural language inference, and question answering.In this paper, we raise concerns that model editing's improvements on factuality may come at the cost of a significant degradation of the model's general abilities.We systematically analyze the side effects by evaluating four popular editing methods on three LLMs across eight representative tasks.Our extensive empirical experiments show that it is challenging for current editing methods to simultaneously improve factuality of LLMs and maintain their general abilities.Our analysis reveals that the side effects are caused by model editing altering the original model weights excessively, leading to overfitting to the edited facts.To mitigate this, a method named RECT is proposed to regularize the edit update weights by imposing constraints on their complexity based on the RElative Change in weighT.Evaluation results show that RECT can significantly mitigate the side effects of editing while still maintaining over 94% editing performance 1 . Jia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu, Zhen-Hua Ling, Kai-Wei Chang 0001, Nanyun Peng 0001 |
EMNLP | 1 |
| 2024 | Leveraging Large Language Models for NLG Evaluation: Advances and ChallengesabstractIn the rapidly evolving domain of Natural Language Generation (NLG) evaluation, introducing Large Language Models (LLMs) has opened new avenues for assessing generated content quality, e.g., coherence, creativity, and context relevance.This paper aims to provide a thorough overview of leveraging LLMs for NLG evaluation, a burgeoning area that lacks a systematic analysis.We propose a coherent taxonomy for organizing existing LLM-based evaluation metrics, offering a structured framework to understand and compare these methods.Our detailed exploration includes critically assessing various LLM-based methodologies, as well as comparing their strengths and limitations in evaluating NLG outputs.By discussing unresolved challenges, including bias, robustness, domain-specificity, and unified evaluation, this paper seeks to offer insights to researchers and advocate for fairer and more advanced NLG evaluation techniques. Zhen Li 0048, Tao Shen 0001, Can Xu 0002, Jia-Chen Gu, Yuxuan Lai, Chongyang Tao, Shuai Ma 0001 |
EMNLP | 5 |
| 2024 | Synchronous Faithfulness Monitoring for Trustworthy Retrieval-Augmented GenerationabstractRetrieval-augmented language models (RALMs) have shown strong performance and wide applicability in knowledge-intensive tasks.However, there are significant trustworthiness concerns as RALMs are prone to generating unfaithful outputs, including baseless information or contradictions with the retrieved context.This paper proposes SYNCHECK, a lightweight monitor that leverages fine-grained decoding dynamics including sequence likelihood, uncertainty quantification, context influence, and semantic alignment to synchronously detect unfaithful sentences.By integrating efficiently measurable and complementary signals, SYNCHECK enables accurate and immediate feedback and intervention, achieving 0.85 AUROC in detecting faithfulness errors across six long-form retrieval-augmented generation tasks, improving prior best method by 4%.Leveraging SYNCHECK, we further introduce FOD, a faithfulness-oriented decoding algorithm guided by beam search for long-form retrieval-augmented generation.Empirical results demonstrate that FOD outperforms traditional strategies such as abstention, reranking, or contrastive decoding significantly in terms of faithfulness, achieving over 10% improvement across six datasets. Di Wu 0054, Jia-Chen Gu, Fan Yin, Nanyun Peng 0001, Kai-Wei Chang 0001 |
EMNLP | 2 |
| 2024 | Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text RetrievalabstractAudio-text retrieval (ATR), which retrieves a relevant caption given an audio clip (A2T) and vice versa (T2A), has recently attracted much research attention. Existing methods typically aggregate information from each modality into a single vector for matching, but this sacrifices local details and can hardly capture intricate relationships within and between modalities. Furthermore, current ATR datasets lack comprehensive alignment information, and simple binary contrastive learning labels overlook the measurement of fine-grained semantic differences between samples. To counter these challenges, we present a novel ATR framework that comprehensively captures the matching relationships of multimodal information from different perspectives and finer granularities. Specifically, a fine-grained alignment method is introduced, achieving a more detail-oriented matching through a multiscale process from local to global levels to capture meticulous cross-modal relationships. In addition, we pioneer the application of cross-modal similarity consistency, leveraging intra-modal similarity relationships as soft supervision to boost more intricate alignment. Extensive experiments validate the effectiveness of our approach, outperforming previous methods by significant margins of at least 3.9% (T2A) / 6.9% (A2T) R@1 on the AudioCaps dataset and 2.9% (T2A) / 5.4% (A2T) R@1 on the Clotho dataset. Jia-Chen Gu, Zhen-Hua Ling |
ICASSP | 2 |
| 2024 | Neighboring Perturbations of Knowledge Editing on Large Language ModelsabstractDespite their exceptional capabilities, large language models (LLMs) are prone to generating unintended text due to false or outdated knowledge. Given the resource-intensive nature of retraining LLMs, there has been a notable increase in the development of knowledge editing. However, current approaches and evaluations rarely explore the perturbation of editing on neighboring knowledge. This paper studies whether updating new knowledge to LLMs perturbs the neighboring knowledge encapsulated within them. Specifically, we seek to figure out whether appending a new answer into an answer list to a factual question leads to catastrophic forgetting of original correct answers in this list, as well as unintentional inclusion of incorrect answers. A metric of additivity is introduced and a benchmark dubbed as Perturbation Evaluation of Appending Knowledge (PEAK) is constructed to evaluate the degree of perturbation to neighboring knowledge when appending new knowledge. Besides, a plug-and-play framework termed Appending via Preservation and Prevention (APP) is proposed to mitigate the neighboring perturbation by maintaining the integrity of the answer list. Experiments demonstrate the effectiveness of APP coupling with four editing methods on three LLMs. Jun-Yu Ma, Zhen-Hua Ling, Ningyu Zhang 0001, Jia-Chen Gu |
ICML | 4 |
| 2024 | Syntax-Augmented Hierarchical Interactive Encoder for Zero-Shot Cross-Lingual Information ExtractionabstractZero-shot cross-lingual information extraction (IE) aims at constructing an IE model for some low-resource target languages, given annotations exclusively in some rich-resource languages. Recent studies have shown language-universal features can bridge the gap between languages. However, prior work has neither explored the potential of establishing interactions between language-universal features and contextual representations nor incorporated features that can effectively model constituent span attributes and relationships between multiple spans. In this study, asyntax-augmentedhierarchicalinteractiveencoder (SHINE) is proposed to transfer cross-lingual IE knowledge. The proposed encoder is capable of interactively capturing complementary information between features and contextual information, to derive language-agnostic representations for various cross-lingual IE tasks. Concretely, a multi-level interaction network is designed to hierarchically interact the complementary information to strengthen domain adaptability. Besides, in addition to the well-studied word-level syntax features of part-of-speech and dependency relation, a new span-level syntax feature of constituency structure is introduced to model the constituent span information which is crucial for IE. Experiments across seven languages on three IE tasks and four benchmarks verify the effectiveness and generalization ability of the proposed method. Jun-Yu Ma, Jia-Chen Gu, Zhen-Hua Ling, Quan Liu 0003, Cong Liu 0006 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | GIFT: Graph-Induced Fine-Tuning for Multi-Party Conversation UnderstandingabstractAddressing the issues of who saying what to whom in multi-party conversations (MPCs) has recently attracted a lot of research attention.However, existing methods on MPC understanding typically embed interlocutors and utterances into sequential information flows, or utilize only the superficial of inherent graph structures in MPCs.To this end, we present a plug-and-play and lightweight method named graph-induced fine-tuning (GIFT) which can adapt various Transformer-based pre-trained language models (PLMs) for universal MPC understanding.In detail, the full and equivalent connections among utterances in regular Transformer ignore the sparse but distinctive dependency of an utterance on another in MPCs.To distinguish different relationships between utterances, four types of edges are designed to integrate graph-induced signals into attention mechanisms to refine PLMs originally designed for processing sequential texts.We evaluate GIFT by implementing it into three PLMs, and test the performance on three downstream tasks including addressee recognition, speaker identification and response selection.Experimental results show that GIFT can significantly improve the performance of three PLMs on three downstream tasks and two benchmarks with only 4 additional parameters per encoding layer, achieving new state-of-theart performance on MPC understanding. Jia-Chen Gu, Zhen-Hua Ling, Quan Liu 0003, Cong Liu 0006 |
ACL (1) | 1 |
| 2023 | MADNet: Maximizing Addressee Deduction Expectation for Multi-Party Conversation GenerationabstractModeling multi-party conversations (MPCs) with graph neural networks has been proven effective at capturing complicated and graphical information flows.However, existing methods rely heavily on the necessary addressee labels and can only be applied to an ideal setting where each utterance must be tagged with an "@" or other equivalent addressee label.To study the scarcity of addressee labels which is a common issue in MPCs, we propose MADNet that maximizes addressee deduction expectation in heterogeneous graph neural networks for MPC generation.Given an MPC with a few addressee labels missing, existing methods fail to build a consecutively connected conversation graph, but only a few separate conversation fragments instead.To ensure message passing between these conversation fragments, four additional types of latent edges are designed to complete a fully-connected graph.Besides, to optimize the edge-typedependent message passing for those utterances without addressee labels, an Expectation-Maximization-based method that iteratively generates silver addressee labels (E step), and optimizes the quality of generated responses (M step), is designed.Experimental results on two Ubuntu IRC channel benchmarks show that MADNet outperforms various baseline models on the task of MPC generation, especially under the more common and challenging setting where part of addressee labels are missing. Jia-Chen Gu, Chao-Hong Tan, Caiyuan Chu, Zhen-Hua Ling, Chongyang Tao, Quan Liu 0003, Cong Liu 0006 |
EMNLP | 1 |
| 2022 | HeterMPC: A Heterogeneous Graph Neural Network for Response Generation in Multi-Party ConversationsabstractJia-Chen Gu, Chao-Hong Tan, Chongyang Tao, Zhen-Hua Ling, Huang Hu, Xiubo Geng, Daxin Jiang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Jia-Chen Gu, Chao-Hong Tan, Chongyang Tao, Zhen-Hua Ling, Huang Hu, Xiubo Geng, Daxin Jiang |
ACL (1) | 1 |
| 2022 | Wider & Closer: Mixture of Short-channel Distillers for Zero-shot Cross-lingual Named Entity RecognitionabstractZero-shot cross-lingual named entity recognition (NER) aims at transferring knowledge from annotated and rich-resource data in source languages to unlabeled and lean-resource data in target languages.Existing mainstream methods based on the teacher-student distillation framework ignore the rich and complementary information lying in the intermediate layers of pre-trained language models, and domaininvariant information is easily lost during transfer.In this study, a mixture of short-channel distillers (MSD) method is proposed to fully interact the rich hierarchical information in the teacher model and to transfer knowledge to the student model sufficiently and efficiently.Concretely, a multi-channel distillation framework is designed for sufficient information transfer by aggregating multiple distillers as a mixture.Besides, an unsupervised method adopting parallel domain adaptation is proposed to shorten the channels between the teacher and student models to preserve domaininvariant features.Experiments on four datasets across nine languages demonstrate that the proposed method achieves new state-of-the-art performance on zero-shot cross-lingual NER and shows great generalization and compatibility across languages and fields. Jun-Yu Ma, Beiduo Chen, Jia-Chen Gu, Zhen-Hua Ling, Wu Guo, Quan Liu 0003, Zhigang Chen 0003, Cong Liu 0006 |
EMNLP | 3 |
| 2022 | Who Says What to Whom: A Survey of Multi-Party ConversationsabstractMulti-party conversations (MPCs) are a more practical and challenging scenario involving more than two interlocutors. This research topic has drawn significant attention from both academia and industry, and it is nowadays counted as one of the most promising research areas in the field of dialogue systems. In general, MPC algorithms aim at addressing the issues of Who says What to Whom, specifically, who speaks, say what, and address whom. The complicated interactions between interlocutors, between utterances, and between interlocutors and utterances develop many variant tasks of MPCs worth investigation. In this paper, we present a comprehensive survey of recent advances in text-based MPCs. In particular, we first summarize recent advances on the research of MPC context modeling including dialogue discourse parsing, dialogue flow modeling and self-supervised training for MPCs. Then we review the state-of-the-art models categorized by Who says What to Whom in MPCs. Finally, we highlight the challenges which are not yet well addressed in MPCs and present future research directions. Jia-Chen Gu, Chongyang Tao, Zhen-Hua Ling |
IJCAI | 1 |
| 2021 | MPC-BERT: A Pre-Trained Language Model for Multi-Party Conversation UnderstandingabstractJia-Chen Gu, Chongyang Tao, Zhenhua Ling, Can Xu, Xiubo Geng, Daxin Jiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jia-Chen Gu, Chongyang Tao, Zhen-Hua Ling, Can Xu 0002, Xiubo Geng, Daxin Jiang |
ACL/IJCNLP (1) | 1 |
| 2021 | Detecting Speaker Personas from Conversational TextsabstractPersonas are useful for dialogue response prediction.However, the personas used in current studies are pre-defined and hard to obtain before a conversation.To tackle this issue, we study a new task, named Speaker Persona Detection (SPD), which aims to detect speaker personas based on the plain conversational text.In this task, a best-matched persona is searched out from candidates given the conversational text.This is a many-to-many semantic matching task because both contexts and personas in SPD are composed of multiple sentences.The long-term dependency and the dynamic redundancy among these sentences increase the difficulty of this task.We build a dataset for SPD, dubbed as Persona Match on Persona-Chat (PMPC).Furthermore, we evaluate several baseline models and propose utterance-to-profile (U2P) matching networks for this task.The U2P models operate at a fine granularity which treat both contexts and personas as sets of multiple sequences.Then, each sequence pair is scored and an interpretable overall score is obtained for a context-persona pair through aggregation.Evaluation results show that the U2P models outperform their baseline counterparts significantly. Jia-Chen Gu, Zhen-Hua Ling, Quan Liu 0003, Zhigang Chen 0003, Xiaodan Zhu 0001 |
EMNLP (1) | 1 |
| 2021 | Have You Made a Decision? Where? A Pilot Study on Interpretability of Polarity Analysis Based on Advising Problem
Tianda Li, Jia-Chen Gu, Hui Liu 0033, Quan Liu 0003, Zhen-Hua Ling, Zhiming Su, Xiaodan Zhu 0001 |
ICASSP | 2 |
| 2021 | Partner Matters! An Empirical Study on Fusing Personas for Personalized Response Selection in Retrieval-Based ChatbotsabstractPersona can function as the prior knowledge for maintaining the consistency of dialogue systems. Most of previous studies adopted the self persona in dialogue whose response was about to be selected from a set of candidates or directly generated, but few have noticed the role of partner in dialogue. This paper makes an attempt to thoroughly explore the impact of utilizing personas that describe either self or partner speakers on the task of response selection in retrieval-based chatbots. Four persona fusion strategies are designed, which assume personas interact with contexts or responses in different ways. These strategies are implemented into three representative models for response selection, which are based on the Hierarchical Recurrent Encoder (HRE), Interactive Matching Network (IMN) and Bidirectional Encoder Representations from Transformers (BERT) respectively. Empirical studies on the Persona-Chat dataset show that the partner personas neglected in previous studies can improve the accuracy of response selection in the IMN- and BERT-based models. Besides, our BERT-based model implemented with the context-response-aware persona fusion strategy outperforms previous methods by margins larger than 2.7% on original personas and 4.6% on revised personas in terms of [email protected] (top-1 accuracy), achieving a new state-of-the-art performance on the Persona-Chat dataset. Jia-Chen Gu, Hui Liu 0033, Zhen-Hua Ling, Quan Liu 0003, Zhigang Chen 0003, Xiaodan Zhu 0001 |
SIGIR | 1 |
| 2021 | Deep Contextualized Utterance Representations for Response Selection and Dialogue AnalysisabstractThe NOESIS II challenge, as the Track 2 in the Eighth Dialogue System Technology Challenge (DSTC 8), is the extension of Track 1 in DSTC 7. Three new elements are incorporated into the extended track, i.e., dialogue with multiple participants, dialogue success, and dialogue disentanglement. These are vital for the creation of a deployed task-oriented dialogue system. This track is divided into four subtasks, the first two of which are evaluated in the form of response selection and the last two focus on dialogue analysis. This paper describes our methods developed for these four subtasks, which all employ deep contextualized utterance representations to make models aware of contextual information and to keep the intrinsic property of multi-turn dialogue systems. In the released evaluation results of Track 2 in DSTC 8, our proposed methods ranked fourth in subtask 1, third in subtask 2, and first in subtask 3 and subtask 4 respectively. In addition to the challenge tasks, we also compare our proposed methods with previous ones on public benchmark datasets. Experimental results show that our proposed methods outperform existing ones by large margins and achieve new state-of-the-art performances on multi-turn response selection and dialogue disentanglement. Jia-Chen Gu, Tianda Li, Zhen-Hua Ling, Quan Liu 0003, Zhiming Su, Yu-Ping Ruan, Xiaodan Zhu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Speaker-Aware BERT for Multi-Turn Response Selection in Retrieval-Based ChatbotsabstractIn this paper, we study the problem of employing pre-trained language models for multi-turn response selection in retrieval-based chatbots. A new model, named Speaker-Aware BERT (SA-BERT), is proposed in order to make the model aware of the speaker change information, which is an important and intrinsic property of multi-turn dialogues. Furthermore, a speaker-aware disentanglement strategy is proposed to tackle the entangled dialogues. This strategy selects a small number of most important utterances as the filtered context according to the speakers' information in them. Finally, domain adaptation is performed to incorporate the in-domain knowledge into pre-trained language models. Experiments on five public datasets show that our proposed model outperforms the present models on all metrics by large margins and achieves new state-of-the-art performances for multi-turn response selection. Jia-Chen Gu, Tianda Li, Quan Liu 0003, Zhen-Hua Ling, Zhiming Su, Si Wei, Xiaodan Zhu 0001 |
CIKM | 1 |
| 2020 | End-to-End Transition-Based Online Dialogue DisentanglementabstractDialogue disentanglement aims to separate intermingled messages into detached sessions. The existing research focuses on two-step architectures, in which a model first retrieves the relationships between two messages and then divides the message stream into separate clusters. Almost all existing work puts significant efforts on selecting features for message-pair classification and clustering, while ignoring the semantic coherence within each session. In this paper, we introduce the first end-to- end transition-based model for online dialogue disentanglement. Our model captures the sequential information of each session as the online algorithm proceeds on processing a dialogue. The coherence in a session is hence modeled when messages are sequentially added into their best-matching sessions. Meanwhile, the research field still lacks data for studying end-to-end dialogue disentanglement, so we construct a large-scale dataset by extracting coherent dialogues from online movie scripts. We evaluate our model on both the dataset we developed and the publicly available Ubuntu IRC dataset [Kummerfeld et al., 2019]. The results show that our model significantly outperforms the existing algorithms. Further experiments demonstrate that our model better captures the sequential semantics and obtains more coherent disentangled sessions. Hui Liu 0033, Jia-Chen Gu, Quan Liu 0003, Si Wei, Xiaodan Zhu 0001 |
IJCAI | 3 |
| 2020 | Generating diverse conversation responses by creating and ranking multiple candidates
Yu-Ping Ruan, Zhen-Hua Ling, Xiaodan Zhu 0001, Quan Liu 0003, Jia-Chen Gu |
Comput. Speech Lang. | 5 |
| 2020 | Utterance-to-Utterance Interactive Matching Network for Multi-Turn Response Selection in Retrieval-Based ChatbotsabstractThis article proposes an utterance-to-utterance interactive matching network (U2U-IMN) for multi-turn response selection in retrieval-based chatbots. Different from previous methods following context-to-response matching or utterance-to-response matching frameworks, this model treats both contexts and responses as sequences of utterances when calculating the matching degrees between them. For a context-response pair, the U2U-IMN model first encodes each utterance separately using recurrent and self-attention layers. Then, a global and bidirectional interaction between the context and the response is conducted using the attention mechanism to collect the matching information between them. The distances between context and response utterances are employed as a prior component when calculating the attention weights. Finally, sentence-level aggregation and context-response-level aggregation are executed in turn to obtain the feature vector for matching degree prediction. Experiments on four public datasets showed that our proposed method outperformed baseline methods on all metrics, achieving a new state-of-the-art performance and demonstrating compatibility across domains for multi-turn response selection. Jia-Chen Gu, Zhen-Hua Ling, Quan Liu 0003 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Interactive Matching Network for Multi-Turn Response Selection in Retrieval-Based ChatbotsabstractIn this paper, we propose an interactive matching network (IMN) for the multi-turn response selection task. First, IMN constructs word representations from three aspects to address the challenge of out-of-vocabulary (OOV) words. Second, an attentive hierarchical recurrent encoder (AHRE), which is capable of encoding sentences hierarchically and generating more descriptive representations by aggregating with an attention mechanism, is designed. Finally, the bidirectional interactions between whole multi-turn contexts and response candidates are calculated to derive the matching information between them. Experiments on four public datasets show that IMN outperforms the baseline models on all metrics, achieving a new state-of-the-art performance and demonstrating compatibility across domains for multi-turn response selection. Jia-Chen Gu, Zhen-Hua Ling, Quan Liu 0003 |
CIKM | 1 |
| 2019 | Dually Interactive Matching Network for Personalized Response Selection in Retrieval-Based ChatbotsabstractJia-Chen Gu, Zhen-Hua Ling, Xiaodan Zhu, Quan Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jia-Chen Gu, Zhen-Hua Ling, Xiaodan Zhu 0001, Quan Liu 0003 |
EMNLP/IJCNLP (1) | 1 |
| 2001 | Steady state hierarchical optimizing control for large-scale industrial processes with fuzzy parametersabstractThis paper presents a new method for steady state hierarchical optimizing control of large-scale industrial processes. Several classical steady state coordination mechanisms are applied to the case in which the model coefficients of each subprocess of a large-scale industrial process are replaced by fuzzy numbers. Hence, each subprocess model is converted into a fuzzy form and then the original crisp programming problem with equality and inequality constraints is transformed into the fuzzy programming problem with fuzzy equality and crisp inequality constraints in each local decision unit. The final solutions are obtained by solving the general mathematical programming problem after the fuzzy equality constraints are converted into crisp inequality constraints. The developed method is mainly used to deal with the model-reality difference caused by either the model coefficients of the subprocess not being known accurately or the model slowly varying during normal operation. Three main types of coordination for processes with fuzzy parameters are derived in this paper: interaction balance method (IBM), interaction prediction method (IPM), and mixed method (MM). Simulation results of two examples show that 1) the proposed method can deal with model-reality difference efficiently, 2) the convergence speed of the on-line coordination for fuzzy parameter processes is faster than that of corresponding coordination for crisp parameter processes, and 3) the objective function of real processes can be improved by using the proposed method compared with the classical case. Furthermore, the studies show that the interaction balance method with global feedback (IBMF) based on a double iterative technique for processes with fuzzy parameters is the coordination algorithm that requires the fewest number of on-line iterations so far. Jia-Chen Gu, Bai-Wu Wan |
IEEE Trans. Syst. Man Cybern. Part C | 1 |