Xingxing Zhang 0002

dblp:59/9985-2 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinations
abstract
Wen Luo, Guangyue Peng, Wei Li, Shaohang Wei, Feifan Song, Liang Wang, Nan Yang, Xingxing Zhang, Jing Jin, Furu Wei, Houfeng Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wen Luo 0001, Guangyue Peng, Wei Li 0101, Shaohang Wei, Feifan Song 0001, Liang Wang 0046, Nan Yang 0002, Xingxing Zhang 0002, Furu Wei, Houfeng Wang
ACL (1)8
2025 Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective
abstract
Yiyao Yu, Yuxiang Zhang, Dongdong Zhang, Xiao Liang, Hengyuan Zhang, Xingxing Zhang, Mahmoud Khademi, Hany Hassan Awadalla, Junjie Wang, Yujiu Yang, Furu Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yiyao Yu, Dongdong Zhang 0001, Xingxing Zhang 0002, Mahmoud Khademi, Hany Hassan, Junjie Wang 0011, Yujiu Yang 0001, Furu Wei
ACL (1)6
2025 Self-Boosting Large Language Models with Synthetic Preference Data
abstract
Through alignment with human preferences, Large Language Models (LLMs) have advanced significantly in generating honest, harmless, and helpful responses. However, collecting high-quality preference data is a resource-intensive and creativity-demanding process, especially for the continual improvement of LLMs. We introduce SynPO, a self-boosting paradigm that leverages synthetic preference data for model alignment. SynPO employs an iterative mechanism wherein a self-prompt generator creates diverse prompts, and a response improver refines model responses progressively. This approach trains LLMs to autonomously learn the generative rewards for their own outputs and eliminates the need for large-scale annotation of prompts and human preferences. After four SynPO iterations, Llama3-8B and Mistral-7B show significant enhancements in instruction-following abilities, achieving over 22.1% win rate improvements on AlpacaEval 2.0 and ArenaHard. Simultaneously, SynPO improves the general performance of LLMs on various tasks, validated by a 3.2 to 5.0 average score increase on the well-recognized Open LLM leaderboard.
Qingxiu Dong, Li Dong 0004, Xingxing Zhang 0002, Zhifang Sui, Furu Wei
ICLR3
2025 Think Only When You Need with Large Hybrid-Reasoning Models
abstract
Recent Large Reasoning Models (LRMs) have shown substantially improved reasoning capabilities over traditional Large Language Models (LLMs) by incorporating extended thinking processes prior to producing final responses. However, excessively lengthy thinking introduces substantial overhead in terms of token consumption and latency, which is unnecessary for simple queries. In this work, we introduce Large Hybrid-Reasoning Models (LHRMs), the first kind of model capable of adaptively determining whether to perform reasoning based on the contextual information of user queries. To achieve this, we propose a two-stage training pipeline comprising Hybrid Fine-Tuning (HFT) as a cold start, followed by online reinforcement learning with the proposed Hybrid Group Policy Optimization (HGPO) to implicitly learn to select the appropriate reasoning mode. Furthermore, we introduce a metric called Hybrid Accuracy to quantitatively assess the model’s capability for hybrid reasoning. Extensive experimental results show that LHRMs can adaptively perform hybrid reasoning on queries of varying difficulty and type. It outperforms existing LRMs and LLMs in reasoning and general capabilities while significantly improving efficiency. Together, our work advocates for a reconsideration of the appropriate use of extended reasoning processes and provides a solid starting point for building hybrid reasoning systems.
Lingjie Jiang, Shaohan Huang, Qingxiu Dong, Zewen Chi, Li Dong 0004, Xingxing Zhang 0002, Tengchao Lv, Lei Cui 0001, Furu Wei
NeurIPS7
2024 MathScale: Scaling Instruction Tuning for Mathematical Reasoning
abstract
Large language models (LLMs) have demonstrated remarkable capabilities in problem-solving. However, their proficiency in solving mathematical problems remains inadequate. We propose MathScale, a simple and scalable method to create high-quality mathematical reasoning data using frontier LLMs (e.g., GPT-3.5). Inspired by the cognitive mechanism in human mathematical learning, it first extracts topics and knowledge points from seed math questions and then build a concept graph, which is subsequently used to generate new math questions. MathScale exhibits effective scalability along the size axis of the math dataset that we generate. As a result, we create a mathematical reasoning dataset (MathScaleQA) containing two million math question-answer pairs. To evaluate mathematical reasoning abilities of LLMs comprehensively, we construct MWPBench, a benchmark of Math Word Problems, which is a collection of 9 datasets (including GSM8K and MATH) covering K-12, college, and competition level math problems. We apply MathScaleQA to fine-tune open-source LLMs (e.g., LLaMA-2 and Mistral), resulting in significantly improved capabilities in mathematical reasoning. Evaluated on MWPBench, MathScale-7B achieves state-of-the-art performance across all datasets, surpassing its best peers of equivalent size by 42.8% in micro average accuracy and 43.6% in macro average accuracy, respectively.
Zhengyang Tang, Xingxing Zhang 0002, Benyou Wang, Furu Wei
ICML2
2024 xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token
abstract
This paper introduces xRAG, an innovative context compression method tailored for retrieval-augmented generation. xRAG reinterprets document embeddings in dense retrieval--traditionally used solely for retrieval--as features from the retrieval modality. By employing a modality fusion methodology, xRAG seamlessly integrates these embeddings into the language model representation space, effectively eliminating the need for their textual counterparts and achieving an extreme compression rate. In xRAG, the only trainable component is the modality bridge, while both the retriever and the language model remain frozen. This design choice allows for the reuse of offline-constructed document embeddings and preserves the plug-and-play nature of retrieval augmentation. Experimental results demonstrate that xRAG achieves an average improvement of over 10% across six knowledge-intensive tasks, adaptable to various language model backbones, ranging from a dense 7B model to an 8x7B Mixture of Experts configuration. xRAG not only significantly outperforms previous context compression methods but also matches the performance of uncompressed models on several datasets, while reducing overall FLOPs by a factor of 3.53. Our work pioneers new directions in retrieval-augmented generation from the perspective of multimodality fusion, and we hope it lays the foundation for future efficient and scalable retrieval-augmented systems.
Xin Cheng 0002, Xun Wang 0012, Xingxing Zhang 0002, Tao Ge 0001, Furu Wei, Huishuai Zhang, Dongyan Zhao 0001
NeurIPS3
2022 Sequence Level Contrastive Learning for Text Summarization
abstract
Contrastive learning models have achieved great success in unsupervised visual representation learning, which maximize the similarities between feature representations of different views of the same image, while minimize the similarities between feature representations of views of different images. In text summarization, the output summary is a shorter form of the input document and they have similar meanings. In this paper, we propose a contrastive learning model for supervised abstractive text summarization, where we view a document, its gold summary and its model generated summaries as different views of the same mean representation and maximize the similarities between them during training. We improve over a strong sequence-to-sequence text generation model (i.e., BART) on three different summarization datasets. Human evaluation also shows that our model achieves better faithfulness ratings compared to its counterpart without contrastive objectives. We release our code at https://github.com/xssstory/SeqCo.
Shusheng Xu, Xingxing Zhang 0002, Yi Wu 0013, Furu Wei
AAAI2
2022 Neural Label Search for Zero-Shot Multi-Lingual Extractive Summarization
abstract
In zero-shot multilingual extractive text summarization, a model is typically trained on English summarization dataset and then applied on summarization datasets of other languages.Given English gold summaries and documents, sentence-level labels for extractive summarization are usually generated using heuristics.However, these monolingual labels created on English datasets may not be optimal on datasets of other languages, for that there is the syntactic or semantic discrepancy between different languages.In this way, it is possible to translate the English dataset to other languages and obtain different sets of labels again using heuristics.To fully leverage the information of these different sets of labels, we propose NLSSum (Neural Label Search for Summarization), which jointly learns hierarchical weights for these different sets of labels together with our summarization model.We conduct multilingual zero-shot summarization experiments on MLSUM and WikiLingua datasets, and we achieve state-of-the-art results using both human and automatic evaluations across these two datasets.
Ruipeng Jia, Xingxing Zhang 0002, Yanan Cao 0001, Zheng Lin 0001, Shi Wang 0002, Furu Wei
ACL (1)2
2022 Attention Temperature Matters in Abstractive Summarization Distillation
abstract
Recent progress of abstractive text summarization largely relies on large pre-trained sequence-to-sequence Transformer models, which are computationally expensive.This paper aims to distill these large models into smaller ones for faster inference and with minimal performance loss.Pseudo-labeling based methods are popular in sequence-tosequence model distillation.In this paper, we find simply manipulating attention temperatures in Transformers can make pseudo labels easier to learn for student models.Our experiments on three summarization datasets show our proposed method consistently improves vanilla pseudo-labeling based methods.Further empirical analysis shows that both pseudo labels and summaries produced by our students are shorter and more abstractive.Our code is available at https://github. com/Shengqiang-Zhang/plate.
Shengqiang Zhang, Xingxing Zhang 0002, Hangbo Bao, Furu Wei
ACL (1)2
2020 Unsupervised Fine-tuning for Text Clustering
abstract
Fine-tuning with pre-trained language models (e.g.BERT) has achieved great success in many language understanding tasks in supervised settings (e.g.text classification).However, relatively little work has been focused on applying pre-trained models in unsupervised settings, such as text clustering.In this paper, we propose a novel method to fine-tune pre-trained models unsupervisedly for text clustering, which simultaneously learns text representations and cluster assignments using a clustering oriented loss.Experiments on three text clustering datasets (namely TREC-6, Yelp, and DBpedia) show that our model outperforms the baseline methods and achieves stateof-the-art results.
Shaohan Huang, Furu Wei, Lei Cui 0001, Xingxing Zhang 0002, Ming Zhou 0001
COLING4
2020 Improving the Efficiency of Grammatical Error Correction with Erroneous Span Detection and Correction
abstract
We propose a novel language-independent approach to improve the efficiency for Grammatical Error Correction (GEC) by dividing the task into two subtasks: Erroneous Span Detection (ESD) and Erroneous Span Correction (ESC).ESD identifies grammatically incorrect text spans with an efficient sequence tagging model.Then, ESC leverages a seq2seq model to take the sentence with annotated erroneous spans as input and only outputs the corrected text for these spans.Experiments show our approach performs comparably to conventional seq2seq approaches in both English and Chinese GEC benchmarks with less than 50% time cost for inference.
Mengyun Chen, Tao Ge 0001, Xingxing Zhang 0002, Furu Wei, Ming Zhou 0001
EMNLP (1)3
2020 Pre-training for Abstractive Document Summarization by Reinstating Source Text
abstract
Abstractive document summarization is usually modeled as a sequence-to-sequence (SEQ2SEQ) learning problem.Unfortunately, training large SEQ2SEQ based summarization models on limited supervised summarization data is challenging.This paper presents three sequence-to-sequence pre-training (in shorthand, STEP) objectives which allow us to pre-train a SEQ2SEQ based abstractive summarization model on unlabeled text.The main idea is that, given an input text artificially constructed from a document, a model is pre-trained to reinstate the original document.These objectives include sentence reordering, next sentence generation and masked document generation, which have close relations with the abstractive document summarization task.Experiments on two benchmark summarization datasets (i.e., CNN/DailyMail and New York Times) show that all three objectives can improve performance upon baselines.Compared to models pre-trained on large-scale data (≥160GB), our method, with only 19GB text for pre-training, achieves comparable results, which demonstrates its effectiveness.Code and models are public available at https://github.com/ zoezou2015/abs_pretraining.
Xingxing Zhang 0002, Wei Lu 0011, Furu Wei, Ming Zhou 0001
EMNLP (1)2
2019 Automatic Grammatical Error Correction for Sequence-to-sequence Text Generation: An Empirical Study
abstract
Sequence-to-sequence (seq2seq) models have achieved tremendous success in text generation tasks.However, there is no guarantee that they can always generate sentences without grammatical errors.In this paper, we present a preliminary empirical study on whether and how much automatic grammatical error correction can help improve seq2seq text generation.We conduct experiments across various seq2seq text generation tasks including machine translation, formality style transfer, sentence compression and simplification.Experiments show the state-of-the-art grammatical error correction system can improve the grammaticality of generated text and can bring taskoriented improvements in the tasks where target sentences are in a formal style.
Tao Ge 0001, Xingxing Zhang 0002, Furu Wei, Ming Zhou 0001
ACL (1)2
2019 HIBERT: Document Level Pre-training of Hierarchical Bidirectional Transformers for Document Summarization
abstract
Neural extractive summarization models usually employ a hierarchical encoder for document encoding and they are trained using sentence-level labels, which are created heuristically using rule-based methods.Training the hierarchical encoder with these inaccurate labels is challenging.Inspired by the recent work on pre-training transformer sentence encoders (Devlin et al., 2018), we propose HIBERT (as shorthand for HIerachical Bidirectional Encoder Representations from Transformers) for document encoding and a method to pre-train it using unlabeled data.We apply the pre-trained HIBERT to our summarization model and it outperforms its randomly initialized counterpart by 1.25 ROUGE on the CNN/Dailymail dataset and by 2.0 ROUGE on a version of New York Times dataset.We also achieve the state-of-the-art performance on these two datasets.
Xingxing Zhang 0002, Furu Wei, Ming Zhou 0001
ACL (1)1
2019 Document-Based Question Answering Improves Query-Focused Multi-document Summarization
Weikang Li, Xingxing Zhang 0002, Yunfang Wu, Furu Wei, Ming Zhou 0001
NLPCC (2)2
2018 Neural Latent Extractive Document Summarization
abstract
Extractive summarization models require sentence-level labels, which are usually created heuristically (e.g., with rule-based methods) given that most summarization datasets only have document-summary pairs.Since these labels might be suboptimal, we propose a latent variable extractive model where sentences are viewed as latent variables and sentences with activated variables are used to infer gold summaries.During training the loss comes directly from gold summaries.Experiments on the CNN/Dailymail dataset show that our model improves over a strong extractive baseline trained on heuristically approximated labels and also performs competitively to several recent models.
Xingxing Zhang 0002, Mirella Lapata, Furu Wei, Ming Zhou 0001
EMNLP1
2017 Dependency Parsing as Head Selection
abstract
Conventional graph-based dependency parsers guarantee a tree structure both during training and inference.Instead, we formalize dependency parsing as the problem of independently selecting the head of each word in a sentence.Our model which we call DENSE (as shorthand for Dependency Neural Selection) produces a distribution over possible heads for each word using features obtained from a bidirectional recurrent neural network.Without enforcing structural constraints during training, DENSE generates (at inference time) trees for the overwhelming majority of sentences, while non-tree outputs can be adjusted with a maximum spanning tree algorithm.We evaluate DENSE on four languages (English, Chinese, Czech, and German) with varying degrees of non-projectivity.Despite the simplicity of the approach, our parsers are on par with the state of the art. 1
Xingxing Zhang 0002, Jianpeng Cheng 0001, Mirella Lapata
EACL (1)1
2017 Sentence Simplification with Deep Reinforcement Learning
abstract
Sentence simplification aims to make sentences easier to read and understand.Most recent approaches draw on insights from machine translation to learn simplification rewrites from monolingual corpora of complex and simple sentences.We address the simplification problem with an encoder-decoder model coupled with a deep reinforcement learning framework.Our model, which we call DRESS (as shorthand for Deep REinforcement Sentence Simplification), explores the space of possible simplifications while learning to optimize a reward function that encourages outputs which are simple, fluent, and preserve the meaning of the input.Experiments on three datasets demonstrate that our model outperforms competitive simplification systems. 1
Xingxing Zhang 0002, Mirella Lapata
EMNLP1
2016 On training the recurrent neural network encoder-decoder for large vocabulary end-to-end speech recognition
abstract
Recently, there has been an increasing interest in end-to-end speech recognition using neural networks, with no reliance on hidden Markov models (HMMs) for sequence modelling as in the standard hybrid framework. The recurrent neural network (RNN) encoderdecoder is such a model, performing sequence to sequence mapping without any predefined alignment. This model first transforms the input sequence into a fixed length vector representation, from which the decoder recovers the output sequence. In this paper, we extend our previous work on this model for large vocabulary end-to-end speech recognition. We first present a more effective stochastic gradient decent (SGD) learning rate schedule that can significantly improve the recognition accuracy. We then extend the decoder with long memory by introducing another recurrent layer that performs implicit language modelling. Finally, we demonstrate that using multiple recurrent layers in the encoder can reduce the word error rate. Our experiments were carried out on the Switchboard corpus using a training set of around 300 hours of transcribed audio data, and we have achieved significantly higher recognition accuracy, thereby reduced the gap compared to the hybrid baseline.
Liang Lu 0001, Xingxing Zhang 0002, Steve Renals
ICASSP2
2016 Top-down Tree Long Short-Term Memory Networks
abstract
Long Short-Term Memory (LSTM) networks, a type of recurrent neural network with a more complex computational unit, have been successfully applied to a variety of sequence modeling tasks.In this paper we develop Tree Long Short-Term Memory (TREELSTM), a neural network model based on LSTM, which is designed to predict a tree rather than a linear sequence.TREELSTM defines the probability of a sentence by estimating the generation probability of its dependency tree.At each time step, a node is generated based on the representation of the generated subtree.We further enhance the modeling power of TREELSTM by explicitly representing the correlations between left and right dependents.Application of our model to the MSR sentence completion challenge achieves results beyond the current state of the art.We also report results on dependency parsing reranking achieving competitive performance.
Xingxing Zhang 0002, Liang Lu 0001, Mirella Lapata
HLT-NAACL1
2015 A study of the recurrent neural network encoder-decoder for large vocabulary speech recognition
abstract
Deep neural networks have advanced the state-of-the-art in automatic speech recognition, when combined with hidden Markov models (HMMs). Recently there has been interest in using systems based on recurrent neural networks (RNNs) to perform sequence modelling directly, without the requirement of an HMM superstructure. In this paper, we study the RNN encoder-decoder approach for large vocabulary end-toend speech recognition, whereby an encoder transforms a sequence of acoustic vectors into a sequence of feature representations, from which a decoder recovers a sequence of words. We investigated this approach on the Switchboard corpus using a training set of around 300 hours of transcribed audio data. Without the use of an explicit language model or pronunciation lexicon, we achieved promising recognition accuracy, demonstrating that this approach warrants further investigation. Index Terms: end-to-end speech recognition, deep neural networks, recurrent neural networks, encoder-decoder.
Liang Lu 0001, Xingxing Zhang 0002, Kyunghyun Cho, Steve Renals
INTERSPEECH2
2014 Chinese Poetry Generation with Recurrent Neural Networks
abstract
We propose a model for Chinese poem generation based on recurrent neural networks which we argue is ideally suited to capturing poetic content and form.Our generator jointly performs content selection ("what to say") and surface realization ("how to say") by learning representations of individual characters, and their combinations into one or more lines as well as how these mutually reinforce and constrain each other.Poem lines are generated incrementally by taking into account the entire history of what has been generated so far rather than the limited horizon imposed by the previous line or lexical n-grams.Experimental results show that our model outperforms competitive Chinese poetry generation systems using both automatic and manual evaluation methods.
Xingxing Zhang 0002, Mirella Lapata
EMNLP1