Can Xu 0002

dblp:33/965-2 · DBLP profile ↗
← Back
47ranked-venue papers
2as first author
34since 2021 · last 2026
0000-0002-1949-5715ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 2 first-author · 33 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RubricBench: Aligning Model-Generated Rubrics with Human Standards
abstract
Junyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu, Yidong Ming, Can Xu, Qingfeng Sun, Kai Zheng, Peng Kang, Xue Liu, Chen Ma. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junyi Zhou 0007, Qiyuan Zhang 0001, Yufei Wang 0005, Fuyuan Lyu, Yidong Ming, Can Xu 0002, Qingfeng Sun, Kai Zheng 0001, Xue (Steve) Liu, Chen Ma 0001
ACL (1)6
2025 WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models
abstract
Huawen Feng, Pu Zhao, Qingfeng Sun, Can Xu, Fangkai Yang, Lu Wang, Qianli Ma, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Huawen Feng, Pu Zhao 0004, Qingfeng Sun, Can Xu 0002, Fangkai Yang, Lu Wang 0029, Qianli Ma 0001, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Qi Zhang 0066
ACL (1)4
2025 WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
abstract
Large language models (LLMs), such as GPT-4, have shown remarkable performance in natural language processing (NLP) tasks, including challenging mathematical reasoning. However, most existing open-source models are only pre-trained on large-scale internet data and without math-related optimization. In this paper, we present WizardMath, which enhances the mathematical reasoning abilities of LLMs, by applying our proposed Reinforcement Learning from Evol-Instruct Feedback (RLEIF) method to the domain of math. Through extensive experiments on two mathematical reasoning benchmarks, namely GSM8k and MATH, we reveal the extraordinary capabilities of our model. Remarkably, WizardMath-Mistral 7B surpasses all other open-source LLMs by a substantial margin. Furthermore, WizardMath 70B even outperforms ChatGPT-3.5, Claude Instant, Gemini Pro and Mistral Medium. Additionally, our preliminary exploration highlights the pivotal role of instruction evolution and process supervision in achieving exceptional math performance.
Qingfeng Sun, Can Xu 0002, Pu Zhao 0004, Jian-Guang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, Yansong Tang, Dongmei Zhang 0001
ICLR3
2025 AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
Mengkang Hu, Pu Zhao 0004, Can Xu 0002, Qingfeng Sun, Jian-Guang Lou, Qingwei Lin, Ping Luo 0002, Saravan Rajmohan
KDD (1)3
2025 RetriEVAL: Evaluating Text Generation with Contextualized Lexical Match
abstract
Pre-trained language models have made significant advancements in text generation tasks. Nevertheless, evaluating the generated text with automatic metrics is still challenging. Compared with supervised metrics, unsupervised metrics which are known for generality and robustness, are frequently employed to assess the quality of generated text efficiently. The representative unsupervised metric BERTScore uses pretrained embedding to calculate the word-to-word similarity across all tokens as evaluation scores, which can introduce potential noise due to the inclusion of tokens that do not contribute significantly to the semantics of the text. Furthermore, its heavy reliance on dense embeddings may lead to lower accuracy when evaluating text outside the common contexts represented in the training data, making it less effective in handling uncommon linguistic patterns Additionally, BERTScore treats all tokens with equal importance and lacks the ability to perform meaningful contextual expansion, which can result in less accurate similarity measurements, particularly when dealing with paraphrased or semantically rich text. To address this problem, we propose an unsupervised automatic evaluation metric inspired by the concept of lexical match in information retrieval. Our method leverages contextualized lexical matching to measure exact matches between identical tokens and dynamically matches different tokens based on their contextualized representations. Experiments on SummEval and Topical-Chat demonstrate our proposed RetriEVAL can correlate better with human judgments than previous unsupervised metrics.
Zhen Li 0048, Xinchi Li, Chongyang Tao, Jiazhan Feng, Tao Shen 0001, Can Xu 0002, Hao Wang 0132, Dongyan Zhao 0001, Shuai Ma 0001
WSDM6
2024 Fine-Grained Distillation for Long Document Retrieval
abstract
Long document retrieval aims to fetch query-relevant documents from a large-scale collection, where knowledge distillation has become de facto to improve a retriever by mimicking a heterogeneous yet powerful cross-encoder. However, in contrast to passages or sentences, retrieval on long documents suffers from the \textit{scope hypothesis} that a long document may cover multiple topics. This maximizes their structure heterogeneity and poses a granular-mismatch issue, leading to an inferior distillation efficacy. In this work, we propose a new learning framework, fine-grained distillation (FGD), for long-document retrievers. While preserving the conventional dense retrieval paradigm, it first produces global-consistent representations crossing different fine granularity and then applies multi-granular aligned distillation merely during training. In experiments, we evaluate our framework on two long-document retrieval benchmarks, which show state-of-the-art performance.
Yucheng Zhou 0001, Tao Shen 0001, Xiubo Geng, Chongyang Tao, Jianbing Shen, Guodong Long, Can Xu 0002, Daxin Jiang
AAAI7
2024 Synergistic Interplay between Search and Large Language Models for Information Retrieval
abstract
Jiazhan Feng, Chongyang Tao, Xiubo Geng, Tao Shen, Can Xu, Guodong Long, Dongyan Zhao, Daxin Jiang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Jiazhan Feng, Chongyang Tao, Xiubo Geng, Tao Shen 0001, Can Xu 0002, Guodong Long, Dongyan Zhao 0001, Daxin Jiang
ACL (1)5
2024 Leveraging Large Language Models for NLG Evaluation: Advances and Challenges
abstract
In the rapidly evolving domain of Natural Language Generation (NLG) evaluation, introducing Large Language Models (LLMs) has opened new avenues for assessing generated content quality, e.g., coherence, creativity, and context relevance.This paper aims to provide a thorough overview of leveraging LLMs for NLG evaluation, a burgeoning area that lacks a systematic analysis.We propose a coherent taxonomy for organizing existing LLM-based evaluation metrics, offering a structured framework to understand and compare these methods.Our detailed exploration includes critically assessing various LLM-based methodologies, as well as comparing their strengths and limitations in evaluating NLG outputs.By discussing unresolved challenges, including bias, robustness, domain-specificity, and unified evaluation, this paper seeks to offer insights to researchers and advocate for fairer and more advanced NLG evaluation techniques.
Zhen Li 0048, Tao Shen 0001, Can Xu 0002, Jia-Chen Gu, Yuxuan Lai, Chongyang Tao, Shuai Ma 0001
EMNLP4
2024 Re-Reading Improves Reasoning in Large Language Models
abstract
To enhance the reasoning capabilities of offthe-shelf Large Language Models (LLMs), we introduce a simple, yet general and effective prompting method, RE2, i.e., Re-Reading the question as input.Unlike most thoughteliciting prompting methods, such as Chain-of-Thought (CoT), which aim to elicit the reasoning process in the output, RE2 shifts the focus to the input by processing questions twice, thereby enhancing the understanding process.Consequently, RE2 demonstrates strong generality and compatibility with most thoughteliciting prompting methods, including CoT.Crucially, RE2 facilitates a "bidirectional" encoding in unidirectional decoder-only LLMs because the first pass could provide global information for the second pass.We begin with a preliminary empirical study as the foundation of RE2, illustrating its potential to enable "bidirectional" attention mechanisms.We then evaluate RE2 on extensive reasoning benchmarks across 14 datasets, spanning 112 experiments, to validate its effectiveness and generality.Our findings indicate that, with the exception of a few scenarios on vanilla ChatGPT, RE2 consistently enhances the reasoning performance of LLMs through a simple re-reading strategy.Further analyses reveal RE2's adaptability, showing how it can be effectively integrated with different LLMs, thought-eliciting prompting, and ensemble strategies. 1
Chongyang Tao, Tao Shen 0001, Can Xu 0002, Guodong Long, Jian-Guang Lou, Shuai Ma 0001
EMNLP4
2024 Automatic Instruction Evolving for Large Language Models
abstract
Fine-tuning large pre-trained language models with Evol-Instruct has achieved encouraging results across a wide range of tasks.However, designing effective evolving methods for instruction evolution requires substantial human expertise.This paper proposes Auto Evol-Instruct, an end-to-end framework that evolves instruction datasets using large language models without any human effort.The framework automatically analyzes and summarizes suitable evolutionary strategies for the given instruction data and iteratively improves the evolving method based on issues exposed during the instruction evolution process.Our extensive experiments demonstrate that the best method optimized by Auto Evol-Instruct outperforms human-designed methods on various benchmarks, including MT-Bench, AlpacaEval, GSM8K, and HumanEval.
Can Xu 0002, Yingxiu Zhao, Jian-Guang Lou, Weizhu Chen
EMNLP2
2024 WizardCoder: Empowering Code Large Language Models with Evol-Instruct
abstract
Code Large Language Models (Code LLMs), such as StarCoder, have demonstrated remarkable performance in various code-related tasks. However, different from their counterparts in the general language modeling field, the technique of instruction fine-tuning remains relatively under-researched in this domain. In this paper, we present Code Evol-Instruct, a novel approach that adapts the Evol-Instruct method to the realm of code, enhancing Code LLMs to create novel models, WizardCoder. Through comprehensive experiments on five prominent code generation benchmarks, namely HumanEval, HumanEval+, MBPP, DS-1000, and MultiPL-E, our models showcase outstanding performance. They consistently outperform all other open-source Code LLMs by a significant margin. Remarkably, WizardCoder 15B even surpasses the well-known closed-source LLMs, including Anthropic's Claude and Google's Bard, on the HumanEval and HumanEval+ benchmarks. Additionally, WizardCoder 34B not only achieves a HumanEval score comparable to GPT3.5 (ChatGPT) but also surpasses it on the HumanEval+ benchmark. Furthermore, our preliminary exploration highlights the pivotal role of instruction complexity in achieving exceptional coding performance.
Can Xu 0002, Pu Zhao 0004, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma 0004, Qingwei Lin, Daxin Jiang
ICLR2
2024 WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions
abstract
Training large language models (LLMs) with open-domain instruction following data brings colossal success. However, manually creating such instruction data is very time-consuming and labor-intensive. Moreover, humans may struggle to produce high-complexity instructions. In this paper, we show an avenue for creating large amounts of instruction data with varying levels of complexity using LLM instead of humans. Starting with an initial set of instructions, we use our proposed Evol-Instruct to rewrite them step by step into more complex instructions. Then, we mix all generated instruction data to fine-tune LLaMA. We call the resulting model WizardLM. Both automatic and human evaluations consistently indicate that WizardLM outperforms baselines such as Alpaca (trained from Self-Instruct) and Vicuna (trained from human-created instructions). The experimental results demonstrate that the quality of instruction-following dataset crafted by Evol-Instruct can significantly improve the performance of LLMs.
Can Xu 0002, Qingfeng Sun, Kai Zheng 0021, Xiubo Geng, Pu Zhao 0004, Jiazhan Feng, Chongyang Tao, Qingwei Lin, Daxin Jiang
ICLR1
2024 WizardArena: Post-training Large Language Models via Simulated Offline Chatbot Arena
abstract
Recent work demonstrates that, post-training large language models with open-domain instruction following data have achieved colossal success. Simultaneously, human Chatbot Arena has emerged as one of the most reasonable benchmarks for model evaluation and developmental guidance. However, the processes of manually curating high-quality training data and utilizing online human evaluation platforms are both expensive and limited. To mitigate the manual and temporal costs associated with post-training, this paper introduces a Simulated Chatbot Arena named WizardArena, which is fully based on and powered by open-source LLMs. For evaluation scenario, WizardArena can efficiently predict accurate performance rankings among different models based on offline test set. For training scenario, we simulate arena battles among various state-of-the-art models on a large scale of instruction data, subsequently leveraging the battle results to constantly enhance target model in both the supervised fine-tuning and reinforcement learning . Experimental results demonstrate that our WizardArena aligns closely with the online human arena rankings, and our models trained on offline extensive battle data exhibit significant performance improvements during SFT, DPO, and PPO stages.
Qingfeng Sun, Can Xu 0002, Pu Zhao 0004, Qingwei Lin, Jian-Guang Lou, Shifeng Chen, Yansong Tang, Weizhu Chen
NeurIPS3
2023 MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain Conversation
abstract
Jiazhan Feng, Qingfeng Sun, Can Xu, Pu Zhao, Yaming Yang, Chongyang Tao, Dongyan Zhao, Qingwei Lin. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jiazhan Feng, Qingfeng Sun, Can Xu 0002, Pu Zhao 0004, Yaming Yang 0001, Chongyang Tao, Dongyan Zhao 0001, Qingwei Lin
ACL (1)3
2023 Iterative Proposal Refinement for Weakly-Supervised Video Grounding
abstract
Weakly-Supervised Video Grounding (WSVG) aims to localize events of interest in untrimmed videos with only video-level annotations. To date, most of the state-of-the-art WSVG methods follow a two-stage pipeline, i.e., firstly generating potential temporal proposals and then grounding with these proposal candidates. Despite the recent progress, existing proposal generation methods suffer from two draw-backs: 1) lack of explicit correspondence modeling; and 2) partial coverage of complex events. To this end, we propose a novel IteRative prOposal refiNement network (dubbed as IRON) to gradually distill the prior knowledge into each proposal and encourage proposals with more complete coverage. Specifically, we set up two lightweight distillation branches to uncover the cross-modal correspondence on both the semantic and conceptual levels. Then, an iterative Label Propagation (LP) strategy is devised to prevent the network from focusing excessively on the most discriminative events instead of the whole sentence content. Precisely, during each iteration, the proposal with the minimal distillation loss and its adjacent ones are regarded as the positive samples, which refines proposal confidence scores in a cascaded manner. Extensive experiments and ablation studies on two challenging WSVG datasets have attested to the effectiveness of our IRON. The code will be available at https://github.com/mengcaopku/IRON.
Meng Cao 0002, Fangyun Wei, Can Xu 0002, Xiubo Geng, Long Chen 0016, Can Zhang 0001, Yuexian Zou, Tao Shen 0001, Daxin Jiang
CVPR3
2023 Investigating the Learning Behaviour of In-Context Learning: A Comparison with Supervised Learning
abstract
Large language models (LLMs) have shown remarkable capacity for in-context learning (ICL), where learning a new task from just a few training examples is done without being explicitly pre-trained. However, despite the success of LLMs, there has been little understanding of how ICL learns the knowledge from the given prompts. In this paper, to make progress toward understanding the learning behaviour of ICL, we train the same LLMs with the same demonstration examples via ICL and supervised learning (SL), respectively, and investigate their performance under label perturbations (i.e., noisy labels and label imbalance) on a range of classification tasks. First, via extensive experiments, we find that gold labels have significant impacts on the downstream in-context performance, especially for large language models; however, imbalanced labels matter little to ICL across all model sizes. Second, when comparing with SL, we show empirically that ICL is less sensitive to label perturbations than SL, and ICL gradually attains comparable performance to SL as the model size increases.
Xindi Wang 0001, Yufei Wang 0003, Can Xu 0002, Xiubo Geng, Chongyang Tao, Frank Rudzicz, Robert E. Mercer, Daxin Jiang
ECAI3
2023 LexLIP: Lexicon-Bottlenecked Language-Image Pre-Training for Large-Scale Image-Text Sparse Retrieval
abstract
Image-text retrieval (ITR) aims to retrieve images or texts that match a query originating from the other modality. The conventional dense retrieval paradigm relies on encoding images and texts into dense representations with dual-stream encoders. However, this approach is limited by slow retrieval speeds in large-scale scenarios. To address this issue, we propose a novel sparse retrieval paradigm for ITR that exploits sparse representations in the vocabulary space for images and texts. This paradigm enables us to leverage bag-of-words models and efficient inverted indexes, significantly reducing retrieval latency. A critical gap emerges from representing continuous image data in a sparse vocabulary space. To bridge this gap, we introduce a novel pre-training framework, Lexicon-Bottlenecked Language-Image Pre-Training (LexLIP), that learns importance-aware lexicon representations. By using lexicon-bottlenecked modules between the dual-stream encoders and weakened text decoders, we are able to construct continuous bag-of-words bottlenecks and learn lexicon-importance distributions. Upon pre-training with same-scale data, our LexLIP achieves state-of-the-art performance on two ITR benchmarks, MSCOCO and Flickr30k. Furthermore, in large-scale retrieval scenarios, LexLIP outperforms CLIP with 5.8× faster retrieval speed and 19.1× less index storage memory. Beyond this, LexLIP surpasses CLIP across 8 out of 10 zero-shot image classification tasks.
Pu Zhao 0004, Can Xu 0002, Xiubo Geng, Tao Shen 0001, Chongyang Tao, Jing Ma 0004, Qingwei Lin, Daxin Jiang
ICCV3
2023 LexMAE: Lexicon-Bottlenecked Pretraining for Large-Scale Retrieval
Tao Shen 0001, Xiubo Geng, Chongyang Tao, Can Xu 0002, Xiaolong Huang 0002, Binxing Jiao, Linjun Yang, Daxin Jiang
ICLR4
2023 KnowDA: All-in-One Knowledge Mixture Model for Data Augmentation in Low-Resource NLP
Yufei Wang 0003, Can Xu 0002, Xiubo Geng, Tao Shen 0001, Chongyang Tao, Daxin Jiang
ICLR3
2023 HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval Generalization
Zefeng Cai, Chongyang Tao, Tao Shen 0001, Can Xu 0002, Xiubo Geng, Xin Lin 0001, Liang He 0001, Daxin Jiang
ICLR4
2023 UnifieR: A Unified Retriever for Large-Scale Retrieval
abstract
Large-scale retrieval is to recall relevant documents from a huge collection given a query. It relies on representation learning to embed documents and queries into a common semantic encoding space. According to the encoding space, recent retrieval methods based on pre-trained language models (PLM) can be coarsely categorized into either dense-vector or lexicon-based paradigms. These two paradigms unveil the PLMs' representation capability in different granularities, i.e., global sequence-level compression and local word-level contexts, respectively. Inspired by their complementary global-local contextualization and distinct representing views, we propose a new learning framework, Unifier, which unifies dense-vector and lexicon-based retrieval in one model with a dual-representing capability. Experiments on passage retrieval benchmarks verify its effectiveness in both paradigms. A uni-retrieval scheme is further presented with even better retrieval quality. We lastly evaluate the model on BEIR benchmark to verify its transferability.
Tao Shen 0001, Xiubo Geng, Chongyang Tao, Can Xu 0002, Guodong Long, Kai Zhang 0033, Daxin Jiang
KDD4
2023 LED: Lexicon-Enlightened Dense Retriever for Large-Scale Retrieval
abstract
Retrieval models based on dense representations in semantic space have become an indispensable branch for first-stage retrieval. These retrievers benefit from surging advances in representation learning towards compressive global sequence-level embeddings. However, they are prone to overlook local salient phrases and entity mentions in texts, which usually play pivot roles in first-stage retrieval. To mitigate this weakness, we propose to make a dense retriever align a well-performing lexicon-aware representation model. The alignment is achieved by weakened knowledge distillations to enlighten the retriever via two aspects – 1) a lexicon-augmented contrastive objective to challenge the dense encoder and 2) a pair-wise rank-consistent regularization to make the dense model’s behavior incline to the other. We evaluate our model on three public benchmarks, which shows that with a comparable lexicon-aware retriever as the teacher, our proposed dense one can bring consistent and significant improvements, and even outdo its teacher. In addition, we show our lexicon-aware distillation strategies are compatible with the standard ranker distillation, which can further lift state-of-the-art performance.1
Kai Zhang 0033, Chongyang Tao, Tao Shen 0001, Can Xu 0002, Xiubo Geng, Binxing Jiao, Daxin Jiang
WWW4
2022 PromDA: Prompt-based Data Augmentation for Low-Resource NLU Tasks
abstract
Yufei Wang, Can Xu, Qingfeng Sun, Huang Hu, Chongyang Tao, Xiubo Geng, Daxin Jiang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yufei Wang 0003, Can Xu 0002, Qingfeng Sun, Huang Hu, Chongyang Tao, Xiubo Geng, Daxin Jiang
ACL (1)2
2022 Contextual Fine-to-Coarse Distillation for Coarse-grained Response Selection in Open-Domain Conversations
abstract
Wei Chen, Yeyun Gong, Can Xu, Huang Hu, Bolun Yao, Zhongyu Wei, Zhihao Fan, Xiaowu Hu, Bartuer Zhou, Biao Cheng, Daxin Jiang, Nan Duan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Wei Chen 0088, Yeyun Gong, Can Xu 0002, Huang Hu, Bolun Yao, Zhongyu Wei, Zhihao Fan, Xiaowu Hu, Bartuer Zhou, Biao Cheng, Daxin Jiang, Nan Duan 0001
ACL (1)3
2022 Multimodal Dialogue Response Generation
abstract
Qingfeng Sun, Yujing Wang, Can Xu, Kai Zheng, Yaming Yang, Huang Hu, Fei Xu, Jessica Zhang, Xiubo Geng, Daxin Jiang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Qingfeng Sun, Can Xu 0002, Kai Zheng 0021, Yaming Yang 0001, Huang Hu, Xiubo Geng, Daxin Jiang
ACL (1)3
2022 PCL: Peer-Contrastive Learning with Diverse Augmentations for Unsupervised Sentence Embeddings
abstract
Learning sentence embeddings in an unsupervised manner is fundamental in natural language processing.Recent common practice is to couple pre-trained language models with unsupervised contrastive learning, whose success relies on augmenting a sentence with a semantically-close positive instance to construct contrastive pairs.Nonetheless, existing approaches usually depend on a monoaugmenting strategy, which causes learning shortcuts towards the augmenting biases and thus corrupts the quality of sentence embeddings.A straightforward solution is resorting to more diverse positives from a multiaugmenting strategy, while an open question remains about how to unsupervisedly learn from the diverse positives but with uneven augmenting qualities in the text field.As one answer, we propose a novel Peer-Contrastive Learning (PCL) with diverse augmentations.PCL constructs diverse contrastive positives and negatives at the group level for unsupervised sentence embeddings.PCL performs peer-positive contrast as well as peer-network cooperation, which offers an inherent anti-bias ability and an effective way to learn from diverse augmentations.Experiments on STS benchmarks verify the effectiveness of PCL against its competitors in unsupervised sentence embeddings. 1
Qiyu Wu 0001, Chongyang Tao, Tao Shen 0001, Can Xu 0002, Xiubo Geng, Daxin Jiang
EMNLP4
2022 Small Changes Make Big Differences: Improving Multi-turn Response Selection in Dialogue Systems via Fine-Grained Contrastive Learning
abstract
Retrieve-based dialogue response selection aims to find a proper response from a candidate set given a multi-turn context.Pre-trained language models (PLMs) based methods have yielded significant improvements on this task.The sequence representation plays a key role in the learning of matching degree between the dialogue context and the response.However, we observe that different context-response pairs sharing the same context always have a greater similarity in the sequence representations calculated by PLMs, which makes it hard to distinguish positive responses from negative ones.Motivated by this, we propose a novel Fine-Grained Contrastive (FGC) learning method for the response selection task based on PLMs.This FGC learning strategy helps PLMs to generate more distinguishable matching representations of each dialogue at fine grains, and further make better predictions on choosing positive responses.Empirical studies on two benchmark datasets demonstrate that the proposed FGC learning method can generally and significantly improve the model performance of existing PLMbased matching models. 1
Can Xu 0002, Huang Hu, Lei Sha, Yan Zhang 0117, Daxin Jiang
INTERSPEECH2
2022 Stylized Knowledge-Grounded Dialogue Generation via Disentangled Template Rewriting
abstract
Qingfeng Sun, Can Xu, Huang Hu, Yujing Wang, Jian Miao, Xiubo Geng, Yining Chen, Fei Xu, Daxin Jiang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Qingfeng Sun, Can Xu 0002, Huang Hu, Jian Miao, Xiubo Geng, Daxin Jiang
NAACL-HLT2
2022 Unsupervised Cross-Domain Adaptation for Response Selection Using Self-Supervised and Adversarial Training
abstract
Recently, many neural context-response matching models have been developed for retrieval-based dialogue systems. Although existing models achieve impressive performance through learning on a large amount of in-domain parallel dialogue data, they usually perform worse in another new domain. How to transfer a response retrieval model trained in high-resource domains to other low-resource domains is a crucial problem for scalable dialogue systems. To this end, we investigate the unsupervised cross-domain adaptation for response selection when the target domain has no parallel dialogue data. Specifically, we propose a two-stage method to adapt a response selection model to a new domain using self-supervised and adversarial training based on pre-trained language models (PLMs). To efficiently incorporate domain awareness and target-domain knowledge to PLMs, we first design a self-supervised post-training procedure, including domain discrimination (DD) task, target-domain masked language model (MLM) task and target-domain next sentence prediction (NSP) task. Based on this, we further conduct the adversarial fine-tuning to empower the model to match the proper response with extracted domain-shared features as much as possible. Experimental results show that our proposed method achieves consistent and significant improvements on several cross-domain response selection datasets.
Jia Li 0012, Chongyang Tao, Huang Hu, Can Xu 0002, Daxin Jiang
WSDM4
2021 Open Domain Dialogue Generation with Latent Images
abstract
We consider grounding open domain dialogues with images. Existing work assumes that both an image and a textual context are available, but image-grounded dialogues by nature are more difficult to obtain than textual dialogues. Thus, we propose learning a response generation model with both image-grounded dialogues and textual dialogues by assuming that the visual scene information at the time of a conversation can be represented by an image, and trying to recover the latent images of the textual dialogues through text-to-image generation techniques. The likelihood of the two types of dialogues is then formulated by a response generator and an image reconstructor that are learned within a conditional variational auto-encoding framework. Empirical studies are conducted in both image-grounded conversation and text-based conversation. In the first scenario, image-grounded dialogues, especially under a low-resource setting, can be effectively augmented by textual dialogues with latent images; while in the second scenario, latent images can enrich the content of responses and at the same time keep them relevant to contexts.
Ze Yang 0001, Wei Wu 0014, Huang Hu, Can Xu 0002, Wei Wang 0301, Zhoujun Li 0001
AAAI4
2021 MPC-BERT: A Pre-Trained Language Model for Multi-Party Conversation Understanding
abstract
Jia-Chen Gu, Chongyang Tao, Zhenhua Ling, Can Xu, Xiubo Geng, Daxin Jiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jia-Chen Gu, Chongyang Tao, Zhen-Hua Ling, Can Xu 0002, Xiubo Geng, Daxin Jiang
ACL/IJCNLP (1)4
2021 Maria: A Visual Experience Powered Conversational Agent
abstract
Zujie Liang, Huang Hu, Can Xu, Chongyang Tao, Xiubo Geng, Yining Chen, Fan Liang, Daxin Jiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zujie Liang, Huang Hu, Can Xu 0002, Chongyang Tao, Xiubo Geng, Daxin Jiang
ACL/IJCNLP (1)3
2021 Learning Neural Templates for Recommender Dialogue System
abstract
Though recent end-to-end neural models have shown the promising progress on Conversational Recommender System (CRS), two key challenges still remain.First, the recommended items cannot be always incorporated into the generated replies precisely and appropriately.Second, only the items mentioned in the training corpus have a chance to be recommended in the conversation.To tackle these challenges, we introduce a novel framework called NTRD for recommender dialogue system that decouples the dialogue generation from the item recommendation.NTRD has two key components, i.e., response template generator and item selector.The former adopts an encoder-decoder model to generate a response template with slot locations tied to target items, while the latter fills in slot locations with the proper items using a sufficient attention mechanism.Our approach combines the strengths of both classical slot filling approaches (that are generally controllable) and modern neural NLG approaches (that are generally more natural and accurate).Extensive experiments on the benchmark RE-DIAL show our NTRD significantly outperforms the previous state-of-the-art methods.Besides, our approach has the unique advantage to produce novel items that do not appear in the training set of dialogue corpus.
Zujie Liang, Huang Hu, Can Xu 0002, Jian Miao, Yingying He, Xiubo Geng, Daxin Jiang
EMNLP (1)3
2021 Neural Rule-Execution Tracking Machine For Transformer-Based Text Generation
abstract
Sequence-to-Sequence (Seq2Seq) neural text generation models, especially the pre-trained ones (e.g., BART and T5), have exhibited compelling performance on various natural language generation tasks. However, the black-box nature of these models limits their application in tasks where specific rules (e.g., controllable constraints, prior knowledge) need to be executed. Previous works either design specific model structures (e.g., Copy Mechanism corresponding to the rule "the generated output should include certain words in the source input'') or implement specialized inference algorithms (e.g., Constrained Beam Search) to execute particular rules through the text generation. These methods require the careful design case-by-case and are difficult to support multiple rules concurrently. In this paper, we propose a novel module named Neural Rule-Execution Tracking Machine (NRETM) that can be equipped into various transformer-based generators to leverage multiple rules simultaneously to guide the neural generation model for superior generation performance in an unified and scalable way. Extensive experiments on several benchmarks verify the effectiveness of our proposed model in both controllable and general text generation tasks.
Yufei Wang 0003, Can Xu 0002, Huang Hu, Chongyang Tao, Stephen Wan 0001, Mark Dras, Mark Johnson 0001, Daxin Jiang
NeurIPS2
2020 Knowledge-Grounded Dialogue Generation with Pre-trained Language Models
abstract
We study knowledge-grounded dialogue generation with pre-trained language models. To leverage the redundant external knowledge under capacity constraint, we propose equipping response generation defined by a pre-trained language model with a knowledge selection module, and an unsupervised approach to jointly optimizing knowledge selection and response generation with unlabeled dialogues. Empirical results on two benchmarks indicate that our model can significantly outperform state-of-the-art methods in both automatic evaluation and human judgment.
Xueliang Zhao, Wei Wu 0014, Can Xu 0002, Chongyang Tao, Dongyan Zhao 0001, Rui Yan 0001
EMNLP (1)3
2020 Low-Resource Knowledge-Grounded Dialogue Generation
Xueliang Zhao, Wei Wu 0014, Chongyang Tao, Can Xu 0002, Dongyan Zhao 0001, Rui Yan 0001
ICLR4
2020 Zero-Resource Knowledge-Grounded Dialogue Generation
abstract
While neural conversation models have shown great potentials towards generating informative and engaging responses via introducing external knowledge, learning such a model often requires knowledge-grounded dialogues that are difficult to obtain. To overcome the data challenge and reduce the cost of building a knowledge-grounded dialogue system, we explore the problem under a zero-resource setting by assuming no context-knowledge-response triples are needed for training. To this end, we propose representing the knowledge that bridges a context and a response and the way that the knowledge is expressed as latent variables, and devise a variational approach that can effectively estimate a generation model from independent dialogue corpora and knowledge corpora. Evaluation results on three benchmarks of knowledge-grounded dialogue generation indicate that our model can achieve comparable performance with state-of-the-art methods that rely on knowledge-grounded dialogues for training, and exhibits a good generalization ability over different datasets.
Can Xu 0002, Wei Wu 0014, Yufan Zhao, Xueliang Zhao, Chongyang Tao
NeurIPS2
2019 One Time of Interaction May Not Be Enough: Go Deep with an Interaction-over-Interaction Network for Response Selection in Dialogues
abstract
Currently, researchers have paid great attention to retrieval-based dialogues in opendomain.In particular, people study the problem by investigating context-response matching for multi-turn response selection based on publicly recognized benchmark data sets.State-of-the-art methods require a response to interact with each utterance in a context from the beginning, but the interaction is performed in a shallow way.In this work, we let utterance-response interaction go deep by proposing an interaction-over-interaction network (IoI).The model performs matching by stacking multiple interaction blocks in which residual information from one time of interaction initiates the interaction process again.Thus, matching information within an utterance-response pair is extracted from the interaction of the pair in an iterative fashion, and the information flows along the chain of the blocks via representations.Evaluation results on three benchmark data sets indicate that IoI can significantly outperform state-of-theart methods in terms of various matching metrics.Through further analysis, we also unveil how the depth of interaction affects the performance of IoI.
Chongyang Tao, Wei Wu 0014, Can Xu 0002, Wenpeng Hu, Dongyan Zhao 0001, Rui Yan 0001
ACL (1)3
2019 Neural Response Generation with Meta-words
abstract
We present open domain response generation with meta-words.A meta-word is a structured record that describes various attributes of a response, and thus allows us to explicitly model the one-to-many relationship within open domain dialogues and perform response generation in an explainable and controllable manner.To incorporate meta-words into generation, we enhance the sequence-to-sequence architecture with a goal tracking memory network that formalizes meta-word expression as a goal and manages the generation process to achieve the goal with a state memory panel and a state controller.Experimental results on two large-scale datasets indicate that our model can significantly outperform several state-ofthe-art generation models in terms of response relevance, response diversity, accuracy of oneto-many modeling, accuracy of meta-word expression, and human evaluation.
Can Xu 0002, Wei Wu 0014, Chongyang Tao, Huang Hu, Matt Schuerman
ACL (1)1
2019 Low-Resource Response Generation with Template Prior
abstract
Ze Yang, Wei Wu, Jian Yang, Can Xu, Zhoujun Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ze Yang 0001, Wei Wu 0014, Jian Yang 0030, Can Xu 0002, Zhoujun Li 0001
EMNLP/IJCNLP (1)4
2019 Read, Attend and Comment: A Deep Architecture for Automatic News Comment Generation
abstract
Ze Yang, Can Xu, Wei Wu, Zhoujun Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ze Yang 0001, Can Xu 0002, Wei Wu 0014, Zhoujun Li 0001
EMNLP/IJCNLP (1)2
2019 A Document-grounded Matching Network for Response Selection in Retrieval-based Chatbots
abstract
We present a document-grounded matching network (DGMN) for response selection that can power a knowledge-aware retrieval-based chatbot system. The challenges of building such a model lie in how to ground conversation contexts with background documents and how to recognize important information in the documents for matching. To overcome the challenges, DGMN fuses information in a document and a context into representations of each other, and dynamically determines if grounding is necessary and importance of different parts of the document and the context through hierarchical interaction with a response at the matching step. Empirical studies on two public data sets indicate that DGMN can significantly improve upon state-of-the-art methods and at the same time enjoys good interpretability.
Xueliang Zhao, Chongyang Tao, Wei Wu 0014, Can Xu 0002, Dongyan Zhao 0001, Rui Yan 0001
IJCAI4
2019 Multi-Representation Fusion Network for Multi-Turn Response Selection in Retrieval-Based Chatbots
abstract
We consider context-response matching with multiple types of representations for multi-turn response selection in retrieval-based chatbots. The representations encode semantics of contexts and responses on words, n-grams, and sub-sequences of utterances, and capture both short-term and long-term dependencies among words. With such a number of representations in hand, we study how to fuse them in a deep neural architecture for matching and how each of them contributes to matching. To this end, we propose a multi-representation fusion network where the representations can be fused into matching at an early stage, at an intermediate stage, or at the last stage. We empirically compare different representations and fusing strategies on two benchmark data sets. Evaluation results indicate that late fusion is always better than early fusion, and by fusing the representations at the last stage, our model significantly outperforms the existing methods, and achieves new state-of-the-art performance on both data sets. Through a thorough ablation study, we demonstrate the effect of each representation to matching, which sheds light on how to select them in practical systems.
Chongyang Tao, Wei Wu 0014, Can Xu 0002, Wenpeng Hu, Dongyan Zhao 0001, Rui Yan 0001
WSDM3
2019 A Sequential Matching Framework for Multi-Turn Response Selection in Retrieval-Based Chatbots
abstract
We study the problem of response selection for multi-turn conversation in retrieval-based chatbots. The task involves matching a response candidate with a conversation context, the challenges for which include how to recognize important parts of the context, and how to model the relationships among utterances in the context. Existing matching methods may lose important information in contexts as we can interpret them with a unified framework in which contexts are transformed to fixed-length vectors without any interaction with responses before matching. This motivates us to propose a new matching framework that can sufficiently carry important information in contexts to matching and model relationships among utterances at the same time. The new framework, which we call a sequential matching framework (SMF), lets each utterance in a context interact with a response candidate at the first step and transforms the pair to a matching vector. The matching vectors are then accumulated following the order of the utterances in the context with a recurrent neural network (RNN) that models relationships among utterances. Context-response matching is then calculated with the hidden states of the RNN. Under SMF, we propose a sequential convolutional network and sequential attention network and conduct experiments on two public data sets to test their performance. Experiment results show that both models can significantly outperform state-of-the-art matching methods. We also show that the models are interpretable with visualizations that provide us insights on how they capture and leverage important information in contexts for matching.
Yu Wu 0012, Wei Wu 0014, Chen Xing, Can Xu 0002, Zhoujun Li 0001, Ming Zhou 0001
Comput. Linguistics4
2018 Knowledge Enhanced Hybrid Neural Network for Text Matching
abstract
Long text brings a big challenge to neural network based text matching approaches due to their complicated structures. To tackle the challenge, we propose a knowledge enhanced hybrid neural network (KEHNN) that leverages prior knowledge to identify useful information and filter out noise in long text and performs matching from multiple perspectives. The model fuses prior knowledge into word representations by knowledge gates and establishes three matching channels with words, sequential structures of text given by Gated Recurrent Units (GRUs), and knowledge enhanced representations. The three channels are processed by a convolutional neural network to generate high level features for matching, and the features are synthesized as a matching score by a multilayer perceptron. In this paper, we focus on exploring the use of taxonomy knowledge for text matching. Evaluation results from extensive experiments on public data sets of question answering and conversation show that KEHNN can significantly outperform state-of-the-art matching models and particularly improve matching accuracy on pairs with long text.
Yu Wu 0012, Wei Wu 0014, Can Xu 0002, Zhoujun Li 0001
AAAI3
2018 Neural Response Generation With Dynamic Vocabularies
abstract
We study response generation for open domain conversation in chatbots. Existing methods assume that words in responses are generated from an identical vocabulary regardless of their inputs, which not only makes them vulnerable to generic patterns and irrelevant noise, but also causes a high cost in decoding. We propose a dynamic vocabulary sequence-to-sequence (DVS2S) model which allows each input to possess their own vocabulary in decoding. In training, vocabulary construction and response generation are jointly learned by maximizing a lower bound of the true objective with a Monte Carlo sampling method. In inference, the model dynamically allocates a small vocabulary for an input with the word prediction model, and conducts decoding only with the small vocabulary. Because of the dynamic vocabulary mechanism, DVS2S eludes many generic patterns and irrelevant words in generation, and enjoys efficient decoding at the same time. Experimental results on both automatic metrics and human annotations show that DVS2S can significantly outperform state-of-the-art methods in terms of response quality, but only requires 60% decoding time compared to the most efficient baseline.
Yu Wu 0012, Wei Wu 0014, Dejian Yang, Can Xu 0002, Zhoujun Li 0001
AAAI4
2018 Playing 20 Question Game with Policy-Based Reinforcement Learning
abstract
The 20 Questions (Q20) game is a well known game which encourages deductive reasoning and creativity.In the game, the answerer first thinks of an object such as a famous person or a kind of animal.Then the questioner tries to guess the object by asking 20 questions.In a Q20 game system, the user is considered as the answerer while the system itself acts as the questioner which requires a good strategy of question selection to figure out the correct object and win the game.However, the optimal policy of question selection is hard to be derived due to the complexity and volatility of the game environment.In this paper, we propose a novel policy-based Reinforcement Learning (RL) method, which enables the questioner agent to learn the optimal policy of question selection through continuous interactions with users.To facilitate training, we also propose to use a reward network to estimate the more informative reward.Compared to previous methods, our RL method is robust to noisy answers and does not rely on the Knowledge Base of objects.Experimental results show that our RL method clearly outperforms an entropy-based engineering system and has competitive performance in a noisyfree simulation environment.
Huang Hu, Xianchao Wu, Bingfeng Luo, Chongyang Tao, Can Xu 0002, Wei Wu 0014
EMNLP5