Qingfu Zhu

dblp:185/0500 · DBLP profile ↗
← Back
30ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0003-3395-222XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 5 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
abstract
Large language models (LLMs) utilize key-value (KV) cache to store historical information during sequence processing. The size of KV cache grows linearly as the length of the sequence extends, which seriously affects memory usage and decoding efficiency. Current methods for KV cache eviction typically utilize the last window from the pre-filling phase as queries to compute the KV importance scores for eviction. Although this scheme is simple to implement, it tends to overly focus on local information, potentially leading to the neglect or omission of crucial global information. To mitigate this issue, we propose **Judge Q**, a novel training method which incorporates a soft token list. This method only tunes the model’s embedding layer at a low training cost. By concatenating the soft token list at the end of the input sequence, we train these tokens' attention map to the original input sequence to align with that of the actual decoded tokens. In this way, the queries corresponding to the soft tokens can effectively capture global information and better evaluate the importance of the keys and values within the KV cache, thus maintaining decoding quality when KV cache is evicted. Under the same eviction budget, our method exhibits less performance degradation compared to existing eviction approaches. We validate our approach through experiments conducted on models such as Llama-3.1-8B-Instruct and Mistral-7B-Instruct-v0.3, using benchmarks including LongBench, RULER, and Needle-in-a-Haystack. Results indicate an improvement of approximately 1 point on the LongBench and over 3 points on RULER. This proposed methodology can be seamlessly integrated into existing open-source models with minimal training overhead, thereby enhancing performance in KV cache eviction scenarios.
Yuzhuang Xu, Shiyu Ji, Yang Xu 0049, Qingfu Zhu, Wanxiang Che
AAAI6
2026 CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
abstract
Large Language Models (LLMs) with Mixture-of-Experts (MoE) architectures are distinguished by their strong performance scaling with increasing parameters across a wide range of tasks, yet they also suffer from substantial computational and storage overheads. Notably, the performance gains of MoE models do not scale proportionally with the growth in expert parameters. While prior works attempt to reduce parameters via expert-level pruning, merging, or decomposition, they still suffer from challenges in both performance and computational efficiency. In this paper, we address these challenges by introducing micro-expert as a finer-grained compression unit that spans across matrices. We first establish a more fundamental perspective, viewing MoE layers as mixtures of micro-experts, and present CAMERA, a lightweight and training-free framework for identifying micro-expert redundancy. Our analysis uncovers significant variance in micro-expert contributions during decoding. Based on this insight, we further propose CAMERA-P, a structured micro-expert pruning framework, and CAMERA-Q, a mixed-precision quantization idea designed for micro-experts. Extensive experiments on nine downstream tasks show that CAMERA-P consistently outperforms strong baselines under pruning ratios ranging from 20% to 60%. Furthermore, CAMERA-Q achieves superior results under aggressive 2-bit quantization, surpassing existing matrix- and channel-level ideas. Notably, our method enables complete micro-expert analysis of Qwen2-57B-A14B in less than 5 minutes on a single NVIDIA A100-40GB GPU.
Yuzhuang Xu, Xu Han 0007, Yuanchi Zhang, Shiyu Ji, Qingfu Zhu, Wanxiang Che
AAAI7
2026 When Does Language Matter? Multilingual Instructions Reveal Step-wise Language Sensitivity in Vision-Language-Action Models
abstract
Vision-Language-Action (VLA) models have shown strong performance in language-conditioned robotic manipulation, yet their robustness to linguistic variation remains poorly understood. In this work, We present the first systematic multilingual evaluation of VLA models by translating the LIBERO benchmark into ten languages, revealing severe performance degradation under non-English instructions, with success rates dropping by 30–50%. Through fine-grained analysis of task executions, we find that language influence is highly non-uniform across steps: certain steps exhibit strong language dependence and dominate overall task failure, while others are largely language-agnostic. Based on this insight, we propose a step-wise inference-time intervention that aligns representations according to step language sensitivity, substantially improving performance under linguistic variation. Our results indicate that language robustness in VLA models is fundamentally a step-wise control problem, highlighting the importance of temporally structured analysis for reliable embodied agents.
Tianhao Niu, Qingfu Zhu, Wanxiang Che
ACL (1)4
2026 Scaling Laws for Code: A More Data-Hungry Regime
abstract
Xianzhen Luo, Wenzhen Zheng, Qingfu Zhu, Rongyi Zhang, Houyi Li, Siming Huang, YuanTao Fan, Wanxiang Che. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xianzhen Luo, Wenzhen Zheng, Qingfu Zhu, Rongyi Zhang, Houyi Li, Siming Huang, YuanTao Fan, Wanxiang Che
ACL (1)3
2026 Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards
abstract
Tianhao Niu, Ziyu Han, Qiguang Chen, Shiqi Zhou, Baocai Shan, Hengjie Fang, Qingfu Zhu, Wanxiang Che. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tianhao Niu, Ziyu Han, Qiguang Chen, Shiqi Zhou, Baocai Shan, Hengjie Fang, Qingfu Zhu, Wanxiang Che
ACL (1)7
2025 OpenCoder: The Open Cookbook for Top-Tier Code Large Language Models
abstract
Siming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Yang Xu, Jian Yang, Jiaheng Liu, Chenchen Zhang, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, Qiufeng Wang, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang, Yang Gao, Jie Fu, Qian Liu, Houyi Li, Ge Zhang, Yuan Qi, Xu Yinghui, Wei Chu, Zili Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Siming Huang, Tianhao Cheng, Jason Klein Liu, Weidi Xu, Jiaran Hao, Liuyihan Song, Jian Yang 0030, Linzheng Chai, Ruifeng Yuan, Xianzhen Luo, YuanTao Fan, Qingfu Zhu, Zhaoxiang Zhang 0001, Yang Gao 0021, Jie Fu 0001, Qian Liu 0033, Houyi Li, Ge Zhang 0009, Yuan Qi 0001
ACL (1)16
2025 Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling
abstract
Xianzhen Luo, Yixuan Wang, Qingfu Zhu, Zhiming Zhang, Xuanyu Zhang, Qing Yang, Dongliang Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xianzhen Luo, Qingfu Zhu, Qing Yang 0033, Dongliang Xu
ACL (1)3
2025 Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits
abstract
The Myers-Briggs Type Indicator (MBTI) is one of the most influential personality theories reflecting individual differences in thinking, feeling, and behaving. MBTI personality detection has garnered considerable research interest and has evolved significantly over the years. However, this task tends to be overly optimistic, as it currently does not align well with the natural distribution of population personality traits. Specifically, the self-reported labels in existing datasets result in data quality issues and the hard labels fail to capture the full range of population personality distributions. In this paper, we identify the task by constructing MBTIBench, the first manually annotated MBTI personality detection dataset with soft labels, under the guidance of psychologists. Our experimental results confirm that soft labels can provide more benefits to other psychological tasks than hard labels. We highlight the polarized predictions and biases in LLMs as key directions for future research.
Bohan Li 0010, Jiannan Guan, Longxu Dou, Yunlong Feng, Dingzirui Wang, Yang Xu 0049, Enbo Wang, Qiguang Chen, Bichen Wang, Xiao Xu 0005, Libo Qin 0001, Qingfu Zhu, Wanxiang Che
COLING14
2025 MURRE: Multi-Hop Table Retrieval with Removal for Open-Domain Text-to-SQL
abstract
The open-domain text-to-SQL task aims to retrieve question-relevant tables from massive databases and generate SQL. However, the performance of current methods is constrained by single-hop retrieval, and existing multi-hop retrieval of open-domain question answering is not directly applicable due to the tendency to retrieve tables similar to the retrieved ones but irrelevant to the question. Since the questions in text-to-SQL usually contain all required information, while previous multi-hop retrieval supplements the questions with retrieved documents. Therefore, we propose the multi-hop table retrieval with removal (MURRE), which removes previously retrieved information from the question to guide the retriever towards unretrieved relevant tables. Our experiments on two open-domain text-to-SQL datasets demonstrate an average improvement of 5.7% over the previous state-of-the-art results.
Xuanliang Zhang, Dingzirui Wang, Longxu Dou, Qingfu Zhu, Wanxiang Che
COLING4
2025 Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code Generation
abstract
Chart2Code has recently received significant attention in the multimodal community due to its potential to reduce the burden of visualization and promote a more detailed understanding of charts.However, existing Chart2Coderelated training datasets suffer from at least one of the following issues: (1) limited scale, (2) limited type coverage, and ( 3) inadequate complexity.To address these challenges, we seek more diverse sources that better align with real-world user distributions and propose dual data synthesis pipelines: (1) Synthesize based on online plotting code.(2) Synthesize based on the chart images in the academic paper.We create a large-scale Chart2Code training dataset Chart2Code53, including 53 chart types, 130K Chart-code pairs based on the pipeline.Experimental results demonstrate that even with few parameters, the model finetuned on Chart2Code53 achieves state-ofthe-art performance on multiple Chart2Code benchmarks within open-source models 1 .
Tianhao Niu, Yiming Cui 0001, Baoxin Wang, Xiao Xu 0005, Qingfu Zhu, Dayong Wu, Shijin Wang 0001, Wanxiang Che
EMNLP6
2025 Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
abstract
Large language models (LLMs) rely on keyvalue cache (KV cache) to accelerate decoding by reducing redundant computations.However, the KV cache memory usage grows substantially with longer text sequences, posing challenges for efficient deployment.Existing KV cache eviction methods prune tokens using prefilling-stage attention scores, causing inconsistency with actual inference queries, especially under tight memory budgets.In this paper, we propose Lookahead Q-Cache (LAQ), a novel eviction framework that generates lowcost pseudo lookahead queries to better approximate the true decoding-stage queries.By using these lookahead queries as the observation window for importance estimation, LAQ achieves more consistent and accurate KV cache eviction aligned with real inference scenarios.Experimental results on LongBench and Needlein-a-Haystack benchmarks show that LAQ outperforms existing methods across various budget levels, achieving a 1 ∼ 4 point improvement on LongBench under limited cache budget.Moreover, LAQ is complementary to existing approaches and can be flexibly combined to yield further improvements.
Shiyu Ji, Yuzhuang Xu, Yang Xu 0049, Qingfu Zhu, Wanxiang Che
EMNLP6
2025 RoT: Enhancing Table Reasoning with Iterative Row-Wise Traversals
abstract
The table reasoning task, crucial for efficient data acquisition, aims to answer questions based on the given table .Recently, reasoning large language models (RLLMs) with Long Chain-of-Thought (Long CoT) significantly enhance reasoning capabilities, leading to brilliant performance on table reasoning.However, Long CoT suffers from high cost for training and exhibits low reliability due to table content hallucinations.Therefore, we propose Rowof-Thought (ROT), which performs iteratively row-wise table traversal, allowing for reasoning extension and reflection-based refinement at each traversal.Scaling reasoning length by rowwise traversal and leveraging reflection capabilities of LLMs, ROT is training-free.The sequential traversal encourages greater attention to the table, thus reducing hallucinations.Experiments show that ROT, using non-reasoning models, outperforms RLLMs by an average of 4.3%, and achieves state-of-the-art results on WikiTableQuestions and TableBench with comparable models, proving its effectiveness.Also, ROT outperforms Long CoT with fewer reasoning tokens, indicating higher efficiency.
Xuanliang Zhang, Dingzirui Wang, Keyan Xu, Qingfu Zhu, Wanxiang Che
EMNLP4
2025 Advancing Tool-Augmented Large Language Models via Meta-Verification and Reflection Learning
abstract
Empowering large language models (LLMs) with effective tool utilization capabilities is crucial for enabling AI agents to solve complex problems. However, current models face two major limitations: (1) unreliable tool planning and invocation due to low-quality instruction datasets (e.g., widespread hallucinated API calls), and (2) weak tool reflection abilities (over 90% of errors cannot be corrected) resulting from static imitation learning. To address these critical limitations, we propose Tool-MVR, a novel Tool-Augmented LLM that achieves comprehensive System 2 reasoning through two key innovations. Specifically, we first introduce Multi-Agent Meta-Verification (MAMV), a systematic pipeline that rigorously validates APIs, queries, and reasoning trajectories to construct ToolBench-V, a new high-quality instruction dataset that addresses the limitation of unreliable tool planning and invocation. Second, we propose Exploration-based Reflection Learning (EXPLORE), which enhances tool reflection capabilities by leveraging tool feedback through a dynamic "Error → Reflection → Correction" learning paradigm, resulting in our reflection dataset ToolBench-R and addressing the critical weakness in tool reflection. Finally, we obtain Tool-MVR by finetuning open-source LLMs (e.g., Qwen-7B) on both ToolBench-V and ToolBench-R. Our experiments demonstrate that Tool-MVR achieves state-of-the-art performance on StableToolBench, surpassing both ToolLLM (by 23.9%) and GPT-4 (by 15.3%) while reducing API calls by 31.4%, with strong generalization capabilities across unseen tools and scenarios. Additionally, on our proposed RefineToolBench, the first benchmark specifically designed to evaluate tool reflection capabilities. Tool-MVR achieves a 58.9% error correction rate, significantly outperforming ToolLLM's 9.1%.
Zhiyuan Ma 0006, Jiayu Liu 0001, Xianzhen Luo, Zhenya Huang, Qingfu Zhu, Wanxiang Che
KDD (2)5
2025 Stealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring
abstract
Honglin Mu, Han He, Yuxin Zhou, Yunlong Feng, Yang Xu, Libo Qin, Xiaoming Shi, Zeming Liu, Xudong Han, Qi Shi, Qingfu Zhu, Wanxiang Che. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Honglin Mu, Han He, Yunlong Feng, Yang Xu 0049, Libo Qin 0001, Zeming Liu, Qi Shi 0002, Qingfu Zhu, Wanxiang Che
NAACL (Long Papers)11
2025 A survey of table reasoning with large language models
Xuanliang Zhang, Dingzirui Wang, Longxu Dou, Qingfu Zhu, Wanxiang Che
Frontiers Comput. Sci.4
2025 CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
abstract
Abstract Powerful large language models (LLMs) are increasingly expected to be deployed with lower computational costs, enabling their capabilities on resource-constrained devices. Post-training quantization (PTQ) has emerged as a star approach to achieve this ambition, with best methods compressing weights to less than 2 bit on average. In this paper, we propose Channel-Relaxed Vector Quantization (CRVQ), a novel technique that significantly improves the performance of PTQ baselines at the cost of only minimal additional bits. This state-of-the-art extreme compression method achieves its results through two key innovations: (1) carefully selecting and reordering a very small subset of critical weight channels, and (2) leveraging extended codebooks to relax the constraint of critical channels. With our method, we demonstrate a 38.9% improvement over the current strongest sub-2-bit PTQ baseline, enabling nearer lossless 1-bit compression. Furthermore, our approach offers flexible customization of quantization bit-width and performance, providing a wider range of deployment options for diverse hardware platforms. Code and checkpoints are available at https://github.com/xuyuzhuang11/CRVQ.
Yuzhuang Xu, Shiyu Ji, Qingfu Zhu, Wanxiang Che
Trans. Assoc. Comput. Linguistics3
2024 Semantic-Guided Generative Image Augmentation Method with Diffusion Models for Image Classification
abstract
Existing image augmentation methods consist of two categories: perturbation-based methods and generative methods. Perturbation-based methods apply pre-defined perturbations to augment an original image, but only locally vary the image, thus lacking image diversity. In contrast, generative methods bring more image diversity in the augmented images but may not preserve semantic consistency, thus may incorrectly change the essential semantics of the original image. To balance image diversity and semantic consistency in augmented images, we propose SGID, a Semantic-guided Generative Image augmentation method with Diffusion models for image classification. Specifically, SGID employs diffusion models to generate augmented images with good image diversity. More importantly, SGID takes image labels and captions as guidance to maintain semantic consistency between the augmented and original images. Experimental results show that SGID outperforms the best augmentation baseline by 1.72% on ResNet-50 (from scratch), 0.33% on ViT (ImageNet-21k), and 0.14% on CLIP-ViT (LAION-2B). Moreover, SGID can be combined with other image augmentation baselines and further improves the overall performance. We demonstrate the semantic consistency and image diversity of SGID through quantitative human and automated evaluations, as well as qualitative case studies.
Bohan Li 0010, Xiao Xu 0005, Yutai Hou, Yunlong Feng, Xuanliang Zhang, Qingfu Zhu, Wanxiang Che
AAAI8
2024 Exploring Hybrid Question Answering via Program-based Prompting
abstract
Question answering over heterogeneous data requires reasoning over diverse sources of data, which is challenging due to the large scale of information and organic coupling of heterogeneous data.Various approaches have been proposed to address these challenges.One approach involves training specialized retrievers to select relevant information, thereby reducing the input length.Another approach is to transform diverse modalities of data into a single modality, simplifying the task difficulty and enabling more straightforward processing.In this paper, we propose HPROPRO, a novel program-based prompting framework for the hybrid question answering task.HPRO-PRO follows the code generation and execution paradigm.In addition, HPROPRO integrates various functions to tackle the hybrid reasoning scenario.Specifically, HPROPRO contains function declaration and function implementation to perform hybrid information-seeking over data from various sources and modalities, which enables reasoning over such data without training specialized retrievers or performing modal transformations.Experimental results on two typical hybrid question answering benchmarks HybridQA and MultiModalQA demonstrate the effectiveness of HPROPRO: it surpasses all baseline systems and achieves the best performances in the few-shot settings on both datasets 1 .
Qi Shi 0002, Qingfu Zhu, Wanxiang Che, Ting Liu 0001
ACL (1)4
2024 Enhancing Numerical Reasoning with the Guidance of Reliable Reasoning Processes
abstract
Numerical reasoning is an essential ability for NLP systems to handle numeric information.Recent research indicates that fine-tuning a small-scale model to learn generating reasoning processes alongside answers can significantly enhance performance.However, current methods have the limitation that most methods generate reasoning processes with large language models (LLMs), which are "unreliable" since such processes could contain information unrelated to the answer.To address this limitation, we introduce Enhancing NumeriCal reasOning with Reliable procEsses (ENCORE), which derives the reliable reasoning process by decomposing the answer formula, ensuring which fully supports the answer.Nevertheless, models could lack enough data to learn the reasoning process generation adequately, since our method generates only one single reasoning process for one formula.To overcome this difficulty, we present a series of pre-training tasks to help models learn the reasoning process generation with synthesized data.The experiments show that ENCORE yields improvement on all five experimental datasets with an average of 1.8%, proving the effectiveness of our method 1 .* Corresponding author. 1 Our code is released in link. 2 For the sake of conciseness in this paper, we collectively refer to these elements as formulas.
Dingzirui Wang, Longxu Dou, Xuanliang Zhang, Qingfu Zhu, Wanxiang Che
ACL (1)4
2024 A Survey on Natural Language Processing for Programming
abstract
Natural language processing for programming aims to use NLP techniques to assist programming. It is increasingly prevalent for its effectiveness in improving productivity. Distinct from natural language, a programming language is highly structured and functional. Constructing a structure-based representation and a functionality-oriented algorithm is at the heart of program understanding and generation. In this paper, we conduct a systematic review covering tasks, datasets, evaluation methods, techniques, and models from the perspective of the structure-based and functionality-oriented property, aiming to understand the role of the two properties in each component. Based on the analysis, we illustrate unexplored areas and suggest potential directions for future work.
Qingfu Zhu, Xianzhen Luo, Fang Liu 0032, Wanxiang Che
LREC/COLING1
2024 Python is Not Always the Best Choice: Embracing Multilingual Program of Thoughts
abstract
Program of Thoughts (PoT) is an approach characterized by its executable intermediate steps, which ensure the accuracy of the logical calculations in the reasoning process.Currently, PoT primarily uses Python.However, relying solely on a single language may result in suboptimal solutions and overlook the potential benefits of other programming languages.In this paper, we conduct comprehensive experiments on the programming languages used in PoT and find that no single language consistently delivers optimal performance across all tasks and models.The effectiveness of each language varies depending on the specific scenarios.Inspired by this, we propose a task and model agnostic approach called MultiPoT, which harnesses strength and diversity from various languages.Experimental results reveal that it significantly outperforms Python Self-Consistency.Furthermore, it achieves comparable or superior performance compared to the best monolingual PoT in almost all tasks across all models.In particular, MultiPoT achieves more than 4.6% improvement on average on ChatGPT (gpt-3.5-turbo-0701) 1 .
Xianzhen Luo, Qingfu Zhu, Libo Qin 0001, Qing Yang 0033, Dongliang Xu, Wanxiang Che
EMNLP2
2024 Make Some Noise: Unlocking Language Model Parallel Inference Capability through Noisy Training
abstract
Yixuan Wang, Xianzhen Luo, Fuxuan Wei, Yijun Liu, Qingfu Zhu, Xuanyu Zhang, Qing Yang, Dongliang Xu, Wanxiang Che. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Xianzhen Luo, Fuxuan Wei, Qingfu Zhu, Qing Yang 0033, Dongliang Xu, Wanxiang Che
EMNLP5
2024 OneBit: Towards Extremely Low-bit Large Language Models
abstract
Model quantification uses low bit-width values to represent the weight matrices of existing models to be quantized, which is a promising approach to reduce both storage and computational overheads of deploying highly anticipated LLMs. However, current quantization methods suffer severe performance degradation when the bit-width is extremely reduced, and thus focus on utilizing 4-bit or 8-bit values to quantize models. This paper boldly quantizes the weight matrices of LLMs to 1-bit, paving the way for the extremely low bit-width deployment of LLMs. For this target, we introduce a 1-bit model compressing framework named OneBit, including a novel 1-bit parameter representation method to better quantize LLMs as well as an effective parameter initialization method based on matrix decomposition to improve the convergence speed of the quantization framework. Sufficient experimental results indicate that OneBit achieves good performance (at least 81% of the non-quantized performance on LLaMA models) with robust training processes when only using 1-bit weight matrices.
Yuzhuang Xu, Xu Han 0007, Zonghan Yang, Shuo Wang 0013, Qingfu Zhu, Zhiyuan Liu 0001, Wanxiang Che
NeurIPS5
2023 A Static and Dynamic Attention Framework for Multi Turn Dialogue Generation
abstract
Recently, research on open domain dialogue systems have attracted extensive interests of academic and industrial researchers. The goal of an open domain dialogue system is to imitate humans in conversations. Previous works on single turn conversation generation have greatly promoted the research of open domain dialogue systems. However, understanding multiple single turn conversations is not equal to the understanding of multi turn dialogue due to the coherent and context dependent properties of human dialogue. Therefore, in open domain multi turn dialogue generation, it is essential to modeling the contextual semantics of the dialogue history rather than only according to the last utterance. Previous research had verified the effectiveness of the hierarchical recurrent encoder-decoder framework on open domain multi turn dialogue generation. However, using an RNN-based model to hierarchically encoding the utterances to obtain the representation of dialogue history still face the problem of a vanishing gradient. To address this issue, in this article, we proposed a static and dynamic attention-based approach to model the dialogue history and then generate open domain multi turn dialogue responses. Experimental results on the Ubuntu and Opensubtitles datasets verify the effectiveness of the proposed static and dynamic attention-based approach on automatic and human evaluation metrics in various experimental settings. Meanwhile, we also empirically verify the performance of combining the static and dynamic attentions on open domain multi turn dialogue generation.
Weinan Zhang 0003, Yiming Cui 0001, Yifa Wang, Qingfu Zhu, Lingzhi Li 0003, Ting Liu 0001
ACM Trans. Inf. Syst.5
2021 Neural Stylistic Response Generation with Disentangled Latent Variables
abstract
Qingfu Zhu, Wei-Nan Zhang, Ting Liu, William Yang Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Qingfu Zhu, Weinan Zhang 0003, Ting Liu 0001, William Yang Wang
ACL/IJCNLP (1)1
2020 Counterfactual Off-Policy Training for Neural Dialogue Generation
abstract
Open-domain dialogue generation suffers from the data insufficiency problem due to the vast size of potential responses.In this paper, we propose to explore potential responses by counterfactual reasoning.Given an observed response, the counterfactual reasoning model automatically infers the outcome of an alternative policy that could have been taken.The resulting counterfactual response synthesized in hindsight is of higher quality than the response synthesized from scratch.Training on the counterfactual responses under the adversarial learning framework helps to explore the high-reward area of the potential response space.An empirical study on the DailyDialog dataset shows that our approach significantly outperforms the HRED model as well as the conventional adversarial learning approaches.
Qingfu Zhu, Weinan Zhang 0003, Ting Liu 0001, William Yang Wang
EMNLP (1)1
2020 Order-Sensitive Keywords Based Response Generation in Open-Domain Conversational Systems
abstract
External keywords are crucial for response generation models to address the generic response problems in open-domain conversational systems. The occurrence of keywords in a response depends heavily on the order of the keywords as they are generated sequentially. Meanwhile, the order of keywords also affects the semantics of a response. Previous keywords based methods mainly focus on the composite of keywords, while the order of keywords has not been sufficiently discussed. In this work, we propose an order-sensitive keywords based model to explore the influence of the order of keywords in open-domain response generation. It automatically inferences the most suitable order that is optimized to generate a natural and relevant response, and subsequently generates the response using the ordered keywords as building blocks. We conducted experiments on a public Twitter dataset and the results show that our approach outperforms the state-of-the-art baselines in both automatic and human evaluations.
Qingfu Zhu, Weinan Zhang 0003, Lei Cui 0003, Ting Liu 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2019 Retrieval-Enhanced Adversarial Training for Neural Response Generation
abstract
Dialogue systems are usually built on either generation-based or retrieval-based approaches, yet they do not benefit from the advantages of different models.In this paper, we propose a Retrieval-Enhanced Adversarial Training (REAT) method for neural response generation.Distinct from existing approaches, the REAT method leverages an encoder-decoder framework in terms of an adversarial training paradigm, while taking advantage of N-best response candidates from a retrieval-based system to construct the discriminator.An empirical study on a large scale public available benchmark dataset shows that the REAT method significantly outperforms the vanilla Seq2Seq model as well as the conventional adversarial training approach.
Qingfu Zhu, Lei Cui 0001, Weinan Zhang 0003, Furu Wei, Ting Liu 0001
ACL (1)1
2019 Neural personalized response generation as domain adaptation
Weinan Zhang 0003, Qingfu Zhu, Yifa Wang, Ting Liu 0001
World Wide Web2
2018 Context-Sensitive Generation of Open-Domain Conversational Responses
abstract
Despite the success of existing works on single-turn conversation generation, taking the coherence in consideration, human conversing is actually a context-sensitive process. Inspired by the existing studies, this paper proposed the static and dynamic attention based approaches for context-sensitive generation of open-domain conversational responses. Experimental results on two public datasets show that the proposed static attention based approach outperforms all the baselines on automatic and human evaluation.
Weinan Zhang 0003, Yiming Cui 0001, Yifa Wang, Qingfu Zhu, Lingzhi Li 0003, Lianqiang Zhou, Ting Liu 0001
COLING4