Hainan Zhang 0001

dblp:120/7359-1 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-7560-1500ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FedSEA-LLaMA: A Secure, Efficient and Adaptive Federated Splitting Framework for Large Language Models
abstract
Private data holds promise for improving LLMs due to its high quality, but its scattered distribution across data silos and the high computational demands of LLMs limit their deployment in federated environments. To address this, the transformer-based federated split models are proposed, which offload most model parameters to the server (or distributed clients) while retaining only a small portion on the client to ensure data privacy. Despite this design, they still face three challenges: 1) Peer-to-peer key encryption struggles to secure transmitted vectors effectively; 2) The auto-regressive nature of LLMs means that federated split learning can only train and infer sequentially, causing high communication overhead; 3) Fixed partition points lack adaptability to downstream tasks. In this paper, we introduce FedSEA-LLaMA, a Secure, Efficient, and Adaptive Federated splitting framework based on LLaMA2. First, we inject Gaussian noise into forward-pass hidden states to enable secure end-to-end vector transmission. Second, we employ attention-mask compression and KV cache collaboration to reduce communication costs, accelerating training and inference. Third, we allow users to dynamically adjust the partition points for input/output blocks based on specific task requirements. Experiments on natural language understanding, summarization, and conversational QA tasks show that FedSEA-LLaMA maintains performance comparable to centralized LLaMA2 and achieves up to 8× speedups in training and inference. Further analysis of privacy attacks and different partition points also demonstrates the effectiveness of FedSEA-LLaMA in security and adaptability.
Zishuai Zhang 0001, Hainan Zhang 0001, Qinnan Zhang, Jin Dong 0004, Yongxin Tong, Zhiming Zheng 0001
AAAI2
2026 Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented Generation
abstract
Retrieval-Augmented Generation (RAG) has become a key paradigm for reducing factual hallucinations in Large Language Models (LLMs), yet little is known about how the order of retrieved documents affects model behavior.We empirically show that under a Top-5 retrieval setting with the gold document included, LLM answers vary substantially across permutations of the retrieved set, even when the gold document is fixed in the first position.This reveals a previously underexplored sensitivity to retrieval permutations.Although existing robust RAG methods focus primarily on enhancing LLM robustness to low-quality retrieval and mitigating positional bias to distribute attention fairly over long contexts, neither approach directly addresses permutation sensitivity.In this paper, we propose Stable-RAG, which exploits permutation sensitivity estimation to mitigate permutation-induced hallucinations.Stable-RAG runs the generator under multiple retrieval orders, clusters hidden states, and decodes from a cluster-center representation that captures the dominant reasoning pattern.It then uses these reasoning results to align hallucinated outputs toward the correct answer, encouraging the model to produce consistent and accurate predictions across document permutations.Experiments on three QA datasets show that Stable-RAG improves answer accuracy, reasoning consistency, and generalization across datasets, retrievers, and input lengths compared with strong baselines 1 .
Qianchi Zhang, Hainan Zhang 0001, Liang Pang 0001, Hongwei Zheng 0003, Zhiming Zheng 0001
ACL (1)2
2026 FedMosaic: Federated Retrieval-Augmented Generation via Parametric Adapters
abstract
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by grounding generation in external knowledge to improve factuality and reduce hallucinations. Yet most deployments assume a centralized corpus, which is infeasible in privacy-aware domains where knowledge remains siloed. This motivates federated RAG (FedRAG), where a central LLM server collaborates with distributed silos without sharing raw documents. In-context RAG violates this requirement by transmitting verbatim documents, whereas parametric RAG encodes documents into light weight adapters that merge with a frozen LLM at inference, avoiding raw-text exchange. We adopt the parametric approach but face two unique challenges induced by FedRAG: high storage and communication from per-document adapters, and destructive aggregation caused by indiscriminately merging multiple adapters. We present FedMosaic, the first federated RAG framework built on parametric adapters. FedMosaic clusters semantically related documents into multi-document adapters with document-specific masks to reduce overhead while preserving specificity, and performs selective adapter aggregation to combine only relevance-aligned, non-conflicting adapters. Experiments show that FedMosaic achieves an average 10.9% higher accuracy than state-of-the-art methods in four categories, while lowering storage costs by 78.8% to 86.3% and communication costs by 91.4%, and never sharing raw documents.
Zhilin Liang, Yuxiang Wang 0014, Zimu Zhou, Hainan Zhang 0001, Boyi Liu 0002, Yongxin Tong
SIGIR4
2026 Less is More: Compact Clue Selection for Efficient Retrieval-Augmented Generation Reasoning
abstract
Current RAG retrievers are designed primarily for human readers, emphasizing complete, readable, and coherent paragraphs. However, Large Language Models (LLMs) benefit more from precise, compact, and well-structured input, which enhances reasoning quality and efficiency. Existing methods rely on reranking or summarization to identify key sentences, but may introduce semantic breaks and unfaithfulness. Thus, efficiently extracting and organizing answer-relevant clues from large-scale documents while reducing LLM reasoning costs remains challenging in RAG systems. Inspired by Occam's razor, we frame LLM-centric retrieval as MinMax optimization: maximizing the extraction of potential clues and reranking them for well-organization, while minimizing reasoning costs by truncating to the smallest sufficient set of clues. In this paper, we propose CompSelect, a compact clue selection mechanism for LLM-centric RAG, consisting of a clue extractor, a reranker, and a truncator. (1) The clue extractor first uses answer-containing sentences as fine-tuning targets, aiming to extract sufficient potential clues; (2) The reranker is trained to prioritize effective clues based on real LLM feedback; (3) The truncator uses the truncated text containing the minimum sufficient clues for answering the question as fine-tuning targets, thereby enabling efficient RAG reasoning. Experiments on three QA datasets demonstrate that CompSelect improves performance while reducing both total and online latency compared to a range of baseline methods. Further analysis also confirms its robustness to unreliable retrieval and generalization across different scenarios.
Qianchi Zhang, Hainan Zhang 0001, Liang Pang 0001, Yongxin Tong, Hongwei Zheng 0003, Zhiming Zheng 0001
WWW2
2026 CodeBC: A more secure large language model for smart contract code generation in blockchain
Lingxiang Wang, Hainan Zhang 0001, Qinnan Zhang, Hongwei Zheng 0003, Jin Dong 0004, Zhiming Zheng 0001
Neurocomputing2
2025 High-Fidelity Polarimetric Implicit 3D Reconstruction with View-Dependent Physical Representation
abstract
Neural implicit methods have made remarkable progress in 3D reconstruction. However, previous methods often assume view-independent properties of target objects, which fails to accurately reconstruct objects with challenging characteristics, such as transparency and high reflectivity. To address this limitation, we propose a polarimetric implicit 3D reconstruction method that integrates geometric and polarization information, enabling the production of high-quality meshes in complex scenes. For high-fidelity surface reconstruction, we introduce a view-dependent physical representation that thoroughly analyzes the subtle physical properties of reflections. The reconstruction process is further enhanced by a simple yet effective view-dependent detection algorithm and optimized using the principles of ray tracing and polarization. Experimental results demonstrate the superior performance of the proposed method in both real and synthetic scenarios.
Sijia Wen, Hainan Zhang 0001, Zhiming Zheng 0001
AAAI3
2025 MaFeRw: Query Rewriting with Multi-Aspect Feedbacks for Retrieval-Augmented Large Language Models
abstract
In a real-world RAG system, the current query often involves spoken ellipses and ambiguous references from dialogue contexts, necessitating query rewriting to better describe user's information needs. However, traditional context-based rewriting has minimal enhancement on downstream generation tasks due to the lengthy process from query rewriting to response generation. Some researchers try to utilize reinforcement learning with generation feedback to assist the rewriter, but this sparse rewards provide little guidance in most cases, leading to unstable training and generation results.We find that user's needs are also reflected in the gold documents, retrieved documents and ground-truth. Therefore, by feeding back these multi-aspect dense rewards to query rewriting, more stable and satisfactory responses can be achieved. In this paper, we propose a novel query rewriting method MaFeRw, which improves RAG performance by integrating multi-aspect feedback from both the retrieval process and generated results. Specifically, we first use manual data to train a T5 model for the rewriter initialization. Next, we design three metrics as reinforcement learning feedback: the similarity between the rewritten query and the gold document, the ranking metrics, and ROUGE between the generation and the ground truth. Inspired by RLAIF, we train three kinds of reward models for the above metrics to achieve more efficient training. Finally, we combine the scores of these reward models as feedback, and use PPO algorithm to explore the optimal query rewriting strategy.Experimental results on two conversational RAG datasets demonstrate that MaFeRw achieves superior generation metrics and more stable training compared to baselines.
Yujing Wang 0010, Hainan Zhang 0001, Liang Pang 0001, Hongwei Zheng 0003, Zhiming Zheng 0001
AAAI2
2025 Defending Against Sophisticated Poisoning Attacks with RL-based Aggregation in Federated Learning
abstract
Federated learning is susceptible to model poisoning attacks, especially those meticulously crafted for servers. Traditional defense methods mainly focus on updating assessments or robust aggregation against manually crafted myopic attacks. When facing advanced attacks, their defense stability is notably insufficient. Therefore, it is imperative to develop adaptive defenses against such advanced poisoning attacks. We find that benign clients exhibit significantly higher data distribution stability than malicious clients in federated learning in both CV and NLP tasks. Therefore, the malicious clients can be recognized by observing the stability of their data distribution. In this paper, we propose AdaAggRL, an RL-based Adaptive Aggregation method, to defend against sophisticated poisoning attacks. Specifically, we first utilize distribution learning to simulate the clients' data distributions. Then, we use maximum mean discrepancy (MMD) to calculate the pairwise similarity of the current local model data distribution, its historical data distribution, and global model data distribution. Finally, we use policy learning to adaptively determine the aggregation weights based on the above similarities. Experiments on four real-world datasets demonstrate that the proposed defense model significantly outperforms widely adopted defense models for sophisticated attacks.
Yujing Wang 0010, Hainan Zhang 0001, Sijia Wen, Wangjie Qiu
AAAI2
2024 Safely Learning with Private Data: A Federated Learning Framework for Large Language Model
abstract
Private data, being larger and quality-higher than public data, can greatly improve large language models (LLM).However, due to privacy concerns, this data is often dispersed in multiple silos, making its secure utilization for LLM training a challenge.Federated learning (FL) is an ideal solution for training models with distributed private data, but traditional frameworks like FedAvg are unsuitable for LLM due to their high computational demands on clients.An alternative, split learning, offloads most training parameters to the server while training embedding and output layers locally, making it more suitable for LLM.Nonetheless, it faces significant challenges in security and efficiency.Firstly, the gradients of embeddings are prone to attacks, leading to potential reverse engineering of private data.Furthermore, the server's limitation of handle only one client's training request at a time hinders parallel training, severely impacting training efficiency.In this paper, we propose a Federated Learning framework for LLM, named FL-GLM, which prevents data leakage caused by both serverside and peer-client attacks while improving training efficiency.Specifically, we first place the input block and output block on local client to prevent embedding gradient attacks from server.Secondly, we employ key-encryption during client-server communication to prevent reverse engineering attacks from peer-clients.Lastly, we employ optimization methods like client-batching or server-hierarchical, adopting different acceleration methods based on the actual computational capabilities of the server.Experimental results on NLU and generation tasks demonstrate that FL-GLM achieves comparable metrics to centralized chatGLM model, validating the effectiveness of our federated learning framework.
Hainan Zhang 0001, Lingxiang Wang, Wangjie Qiu, Hongwei Zheng 0003, Zhi Ming Zheng
EMNLP2
2024 Debiasing Counterfactual Context With Causal Inference for Multi-Turn Dialogue Reasoning
abstract
In the multi-turn dialogue reasoning task, existing models conduct word-level interaction on the entire context to gather reasoning evidence, which aims to select the logically correct one from the candidate response options. Observing the fact that the salient reasoning evidence usually comes from certain snippets of the whole dialogue session, one promising study direction is to explicitly identify the candidate reasoning contexts correlated with the dialogue reasoning options, called option-related contexts, and then make logical inference among them. However, such option-related contexts are stained with noisy information. As a result, existing models may reason unfairly with biased context and select wrong options. To tackle the context bias problem, in this article, we propose a novel CounterFactual learning framework for Dialogue Reasoning, named CF-DialReas, which mitigates the bias information by subtracting the counterfactual representation from the total causal representation. Specifically, we consider two scenarios, i.e., factual dialogue reasoning where the whole context is available to estimate the total causal representation, and the counterfactual dialogue reasoning, which firstly utilizes three different types of utterance selectors to select option-unrelatedcontext, and then only the option-unrelatedcontext is available to guess the counterfactual representation. Experimental results on two public dialogue reasoning datasets show that the model with our mechanism can obtain higher ranking measures, validating the effectiveness of counterfactual learning of CF-DialReas. Further analysis on the generality of CF-DialReas shows that our counterfactual learning mechanism is generally effective to the widely-used models.
Hainan Zhang 0001, Shuai Zhao 0001, Hongshen Chen, Zhuoye Ding, Zhiguo Wan, Bo Cheng 0001, Yanyan Lan
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 HiBERT: Detecting the illogical patterns with hierarchical BERT for multi-turn dialogue reasoning
Hainan Zhang 0001, Shuai Zhao 0001, Hongshen Chen, Bo Cheng 0001, Zhuoye Ding, Sulong Xu, Weipeng Yan, Yanyan Lan
Neurocomputing2
2022 Automatic Product Copywriting for E-commerce
abstract
Product copywriting is a critical component of e-commerce recommendation platforms. It aims to attract users' interest and improve user experience by highlighting product characteristics with textual descriptions. In this paper, we report our experience deploying the proposed Automatic Product Copywriting Generation (APCG) system into the JD.com e-commerce product recommendation platform. It consists of two main components: 1) natural language generation, which is built from a transformer-pointer network and a pre-trained sequence-to-sequence model based on millions of training data from our in-house platform; and 2) copywriting quality control, which is based on both automatic evaluation and human screening. For selected domains, the models are trained and updated daily with the updated training data. In addition, the model is also used as a real-time writing assistant tool on our live broadcast platform. The APCG system has been deployed in JD.com since Feb 2021. By Sep 2021, it has generated 2.53 million product descriptions, and improved the overall averaged click-through rate (CTR) and the Conversion Rate (CVR) by 4.22% and 3.61%, compared to baselines, respectively on a year-on-year basis. The accumulated Gross Merchandise Volume (GMV) made by our system is improved by 213.42%, compared to the number in Feb 2021.
Yanyan Zou 0003, Hainan Zhang 0001, Shiliang Diao, Zhuoye Ding, Xueqi He, Bo Long, Han Yu 0001, Lingfei Wu 0001
AAAI3
2022 From spoken dialogue to formal summary: An utterance rewriting for dialogue summarization
abstract
Yue Fang, Hainan Zhang, Hongshen Chen, Zhuoye Ding, Bo Long, Yanyan Lan, Yanquan Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yue Fang 0001, Hainan Zhang 0001, Hongshen Chen, Zhuoye Ding, Bo Long, Yanyan Lan, Yanquan Zhou
NAACL-HLT2
2021 Probing Product Description Generation via Posterior Distillation
abstract
In product description generation (PDG), the user-cared aspect is critical for the recommendation system, which can not only improve user's experiences but also obtain more clicks. High-quality customer reviews can be considered as an ideal source to mine user-cared aspects. However, in reality, a large number of new products (known as long-tailed commodities) cannot gather sufficient amount of customer reviews, which brings a big challenge in the product description generation task. Existing works tend to generate the product description solely based on item information, i.e., product attributes or title words, which leads to tedious contents and cannot attract customers effectively. To tackle this problem, we propose an adaptive posterior network based on Transformer architecture that can utilize user-cared information from customer reviews. Specifically, we first extend the self-attentive Transformer encoder to encode product titles and attributes. Then, we apply an adaptive posterior distillation module to utilize useful review information, which integrates user-cared aspects to the generation process. Finally, we apply a Transformer-based decoding phase with copy mechanism to automatically generate the product description. Besides, we also collect a large-scare Chinese product description dataset to support our work and further research in this field. Experimental results show that our model is superior to traditional generative models in both automatic indicators and human evaluation.
Haolan Zhan, Hainan Zhang 0001, Hongshen Chen, Lei Shen 0001, Zhuoye Ding, Yongjun Bao, Weipeng Yan, Yanyan Lan
AAAI2
2021 Adaptive Bridge between Training and Inference for Dialogue Generation
abstract
Although exposure bias has been widely studied in some NLP tasks, it faces its unique challenges in dialogue response generation, the representative one-to-various generation scenario.In real human dialogue, there are many appropriate responses for the same context, not only with different expressions, but also with different topics.Therefore, due to the much bigger gap between various ground-truth responses and the generated synthetic response, exposure bias is more challenging in dialogue generation task.What's more, as MLE encourages the model to only learn the common words among different ground-truth responses, but ignores the interesting and specific parts, exposure bias may further lead to the common response generation problem, such as "I don't know" and "HaHa?"In this paper, we propose a novel adaptive switching mechanism, which learns to automatically transit between ground-truth learning and generated learning regarding the word-level matching score, such as the cosine similarity.Experimental results on both Chinese STC dataset and English Reddit dataset, show that our adaptive method achieves a significant improvement in terms of metric-based evaluation and human evaluation, as compared with the state-of-the-art exposure bias approaches.Further analysis on NMT task also shows that our model can achieve a significant improvement.
Hainan Zhang 0001, Yanyan Zou 0003, Hongshen Chen, Zhuoye Ding, Yanyan Lan
EMNLP (1)2
2021 CoLV: A Collaborative Latent Variable Model for Knowledge-Grounded Dialogue Generation
abstract
Knowledge-grounded dialogue generation has achieved promising performance with the engagement of external knowledge sources.Typical approaches towards this task usually perform relatively independent two sub-tasks, i.e., knowledge selection and knowledge-aware response generation.In this paper, in order to improve the diversity of both knowledge selection and knowledge-aware response generation, we propose a collaborative latent variable (CoLV) model to integrate these two aspects simultaneously in separate yet collaborative latent spaces, so as to capture the inherent correlation between knowledge selection and response generation.During generation, our proposed model firstly draws knowledge candidate from the latent space conditioned on the dialogue context, and then samples a response from another collaborative latent space conditioned on both the context and the selected knowledge.Experimental results on two widely-used knowledge-grounded dialogue datasets show that our model outperforms previous methods on both knowledge selection and response generation.
Haolan Zhan, Lei Shen 0001, Hongshen Chen, Hainan Zhang 0001
EMNLP (1)4
2021 Augmenting Knowledge-grounded Conversations with Sequential Knowledge Transition
abstract
Haolan Zhan, Hainan Zhang, Hongshen Chen, Zhuoye Ding, Yongjun Bao, Yanyan Lan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Haolan Zhan, Hainan Zhang 0001, Hongshen Chen, Zhuoye Ding, Yongjun Bao, Yanyan Lan
NAACL-HLT2
2020 Modeling Topical Relevance for Multi-Turn Dialogue Generation
abstract
Topic drift is a common phenomenon in multi-turn dialogue. Therefore, an ideal dialogue generation models should be able to capture the topic information of each context, detect the relevant context, and produce appropriate responses accordingly. However, existing models usually use word or sentence level similarities to detect the relevant contexts, which fail to well capture the topical level relevance. In this paper, we propose a new model, named STAR-BTM, to tackle this problem. Firstly, the Biterm Topic Model is pre-trained on the whole training dataset. Then, the topic level attention weights are computed based on the topic representation of each context. Finally, the attention weights and the topic distribution are utilized in the decoding process to generate the corresponding responses. Experimental results on both Chinese customer services data and English Ubuntu dialogue data show that STAR-BTM significantly outperforms several state-of-the-art methods, in terms of both metric-based and human evaluations.
Hainan Zhang 0001, Yanyan Lan, Liang Pang 0001, Hongshen Chen, Zhuoye Ding, Dawei Yin 0001
IJCAI1
2020 User-Inspired Posterior Network for Recommendation Reason Generation
abstract
Recommendation reason generation, aiming at showing the selling points of products for customers, plays a vital role in attracting customers' attention as well as improving user experience. A simple and effective way is to extract keywords directly from the knowledge-base of products, i.e., attributes or title, as the recommendation reason. However, generating recommendation reason from product knowledge doesn't naturally respond to users' interests. Fortunately, on some E-commerce websites, there exists more and more user-generated content (user-content for short), i.e., product question-answering (QA) discussions, which reflect user-cared aspects. Therefore, in this paper, we consider generating the recommendation reason by taking into account not only the product attributes but also the customer-generated product QA discussions. In reality, adequate user-content is only possible for the most popular commodities, whereas large sums of long-tail products or new products cannot gather a sufficient number of user-content. To tackle this problem, we propose a user-inspired multi-source posterior transformer (MSPT), which induces the model reflecting the users' interests with a posterior multiple QA discussions module, and generating recommendation reasons containing the product attributes as well as the user-cared aspects. Experimental results show that our model is superior to traditional generative models. Additionally, the analysis also shows that our model can focus more on the user-cared aspects than baselines.
Haolan Zhan, Hainan Zhang 0001, Hongshen Chen, Lei Shen 0001, Yanyan Lan, Zhuoye Ding, Dawei Yin 0001
SIGIR2
2019 ReCoSa: Detecting the Relevant Contexts with Self-Attention for Multi-turn Dialogue Generation
abstract
In multi-turn dialogue generation, response is usually related with only a few contexts.Therefore, an ideal model should be able to detect these relevant contexts and produce a suitable response accordingly.However, the widely used hierarchical recurrent encoderdecoder models just treat all the contexts indiscriminately, which may hurt the following response generation process.Some researchers try to use the cosine similarity or the traditional attention mechanism to find the relevant contexts, but they suffer from either insufficient relevance assumption or position bias problem.In this paper, we propose a new model, named ReCoSa, to tackle this problem.Firstly, a word level LSTM encoder is conducted to obtain the initial representation of each context.Then, the self-attention mechanism is utilized to update both the context and masked response representation.Finally, the attention weights between each context and response representations are computed and used in the further decoding process.Experimental results on both Chinese customer services dataset and English Ubuntu dialogue dataset show that ReCoSa significantly outperforms baseline models, in terms of both metric-based and human evaluations.Further analysis on attention shows that the detected relevant contexts by ReCoSa are highly coherent with human's understanding, validating the correctness and interpretability of ReCoSa.
Hainan Zhang 0001, Yanyan Lan, Liang Pang 0001, Jiafeng Guo, Xueqi Cheng 0001
ACL (1)1
2018 Tailored Sequence to Sequence Models to Different Conversation Scenarios
abstract
Sequence to sequence (Seq2Seq) models have been widely used for response generation in the area of conversation.However, the requirements for different conversation scenarios are distinct.For example, customer service requires the generated responses to be specific and accurate, while chatbot prefers diverse responses so as to attract different users.The current Seq2Seq model fails to meet these diverse requirements, by using a general average likelihood as the optimization criteria.As a result, it usually generates safe and commonplace responses, such as 'I don't know'.In this paper, we propose two tailored optimization criteria for Seq2Seq to different conversation scenarios, i.e., the maximum generated likelihood for specific-requirement scenario, and the conditional value-at-risk for diverse-requirement scenario.Experimental results on the Ubuntu dialogue corpus (Ubuntu service scenario) and Chinese Weibo dataset (social chatbot scenario) show that our proposed models not only satisfies diverse requirements for different scenarios, but also yields better performances against traditional Seq2Seq models in terms of both metric-based and human evaluations.
Hainan Zhang 0001, Yanyan Lan, Jiafeng Guo, Jun Xu 0001, Xueqi Cheng 0001
ACL (1)1
2018 Reinforcing Coherence for Sequence to Sequence Model in Dialogue Generation
abstract
Sequence to sequence (Seq2Seq) approach has gained great attention in the field of single-turn dialogue generation. However, one serious problem is that most existing Seq2Seq based models tend to generate common responses lacking specific meanings. Our analysis show that the underlying reason is that Seq2Seq is equivalent to optimizing Kullback–Leibler (KL) divergence, thus does not penalize the case whose generated probability is high while the true probability is low. However, the true probability is unknown, which poses challenges for tackling this problem. Inspired by the fact that the coherence (i.e. similarity) between post and response is consistent with human evaluation, we hypothesize that the true probability of a response is proportional to the coherence degree. The coherence scores are then used as the reward function in a reinforcement learning framework to penalize the case whose generated probability is high while the true probability is low. Three different types of coherence models, including an unlearned similarity function, a pretrained semantic matching function, and an end-to-end dual learning architecture, are proposed in this paper. Experimental results on both Chinese Weibo dataset and English Subtitle dataset show that the proposed models produce more specific and meaningful responses, yielding better performances against Seq2Seq models in terms of both metric-based and human evaluations.
Hainan Zhang 0001, Yanyan Lan, Jiafeng Guo, Jun Xu 0001, Xueqi Cheng 0001
IJCAI1