VLDB 2026 Research / reviewers in the wild / expert
Xing Wu 0002
dblp:04/55-2
· DBLP profile ↗
16ranked-venue papers
5as first author
14since 2021 · last 2026
0009-0004-3796-3705ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMsabstractHigh-quality long-context data is essential for training large language models (LLMs) capable of processing extensive documents, yet existing synthesis approaches using relevance-based aggregation face challenges of computational efficiency. We present LiteLong, a resource-efficient method for synthesizing long-context data through structured topic organization and multi-agent debate. Our approach leverages the BISAC book classification system to provide a comprehensive hierarchical topic organization, and then employs a debate mechanism with multiple LLMs to generate diverse, high-quality topics within this structure. For each topic, we use lightweight BM25 retrieval to obtain relevant documents and concatenate them into 128K-token training samples. Experiments on HELMET and Ruler benchmarks demonstrate that LiteLong achieves competitive long-context performance and can seamlessly integrate with other long-dependency enhancement methods. LiteLong makes high-quality long-context data synthesis more accessible by reducing both computational and data engineering costs, facilitating further research in long-context language training. Junlong Jia, Xing Wu 0002, Chaochen Gao, Zijia Lin, Zhongzhi Li, Weinong Wang, Donghui Jin, Debing Zhang |
AAAI | 2 |
| 2025 | Task-level Distributionally Robust Optimization for Large Language Model-based Dense RetrievalabstractLarge Language Model-based Dense Retrieval (LLM-DR) optimizes over numerous heterogeneous fine-tuning collections from different domains. However, the discussion about its training data distribution is still minimal. Previous studies rely on empirically assigned dataset choices or sampling ratios, which inevitably lead to sub-optimal retrieval performances. In this paper, we propose a new task-level Distributionally Robust Optimization (tDRO) algorithm for LLM-DR fine-tuning, targeted at improving the universal domain generalization ability by end-to-end reweighting the data distribution of each task. The tDRO parameterizes the domain weights and updates them with scaled domain gradients. The optimized weights are then transferred to the LLM-DR fine-tuning to train more robust retrievers. Experiments show optimal improvements in large-scale retrieval benchmarks and reduce up to 30% dataset usage after applying our optimization algorithm with a series of different-sized LLM-DR models. Guangyuan Ma, Yongliang Ma, Xing Wu 0002, Zhenpeng Su, Ming Zhou 0001, Songlin Hu 0001 |
AAAI | 3 |
| 2025 | Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language ModelabstractRecent advancements in large language models (LLMs) have highlighted the importance of extending context lengths for handling complex tasks. While traditional methods for training on long contexts often use filtered long documents, these approaches lead to domain imbalances, limiting model performance. To address this, techniques like random document concatenation (Standard) and similarity-based methods (KNN, ICLM) have been developed. However, they either sacrifice semantic coherence or diversity. To balance both aspects, we introduce Quest, a query-centric data synthesis method aggregating semantically relevant yet diverse documents. Quest uses a generative model to predict potential queries for each document, grouping documents with similar queries and keywords. Extensive experiments demonstrate Quest's superior performance on long-context tasks, achieving remarkable results with context lengths of up to 1M tokens and confirming its scalability across various model sizes. Chaochen Gao, Xing Wu 0002, Songlin Hu 0001 |
ICLR | 2 |
| 2025 | NExtLong: Toward Effective Long-Context Training without Long DocumentsabstractLarge language models (LLMs) with extended context windows have made significant strides yet remain a challenge due to the scarcity of long documents. Existing methods tend to synthesize long-context data but lack a clear mechanism to reinforce the long-range dependency modeling. To address this limitation, we propose NExtLong, a novel framework for synthesizing long-context data through Negative document Extension. NExtLong decomposes a document into multiple meta-chunks and extends the context by interleaving hard negative distractors retrieved from pretraining corpora. This approach compels the model to discriminate long-range dependent context from distracting content, enhancing its ability to model long-range dependencies. Extensive experiments demonstrate that NExtLong achieves significant performance improvements on the HELMET and RULER benchmarks compared to existing long-context synthesis approaches and leading models, which are trained on non-synthetic long documents. These findings highlight NExtLong's ability to reduce reliance on non-synthetic long documents, making it an effective framework for developing advanced long-context LLMs. Chaochen Gao, Xing Wu 0002, Zijia Lin, Debing Zhang, Songlin Hu 0001 |
ICML | 2 |
| 2025 | CartesianMoE: Boosting Knowledge Sharing among Experts via Cartesian Product Routing in Mixture-of-ExpertsabstractZhenpeng Su, Xing W, Zijia Lin, Yizhe Xiong, Minxuan Lv, Guangyuan Ma, Hui Chen, Songlin Hu, Guiguang Ding. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zhenpeng Su, Xing Wu 0002, Zijia Lin, Yizhe Xiong, Minxuan Lv, Guangyuan Ma, Hui Chen 0013, Songlin Hu 0001, Guiguang Ding |
NAACL (Long Papers) | 2 |
| 2025 | LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context InstructionsabstractHigh-quality long-context instruction data is essential for aligning long-context large language models (LLMs). Despite the public release of models like Qwen and Llama, their long-context instruction data remains proprietary. Human annotation is costly and challenging, while template-based synthesis methods limit scale, diversity, and quality. We introduce LongMagpie, a self-synthesis framework that automatically generates large-scale long-context instruction data. Our key insight is that aligned long-context LLMs, when presented with a document followed by special tokens preceding a user turn, auto-regressively generate contextually relevant queries. By harvesting these document-query pairs and the model's responses, LongMagpie produces high-quality instructions without human effort. Experiments on HELMET, RULER, and Longbench v2 demonstrate that LongMagpie achieves leading performance on long-context tasks while maintaining competitive performance on short-context tasks, establishing it as a simple and effective approach for open, diverse, and scalable long-context instruction data synthesis. Chaochen Gao, Xing Wu 0002, Zijia Lin, Debing Zhang, Songlin Hu 0001 |
NeurIPS | 2 |
| 2025 | Libra: Large Chinese-based Safeguard for AI Content
Huimu Yu, Xing Wu 0002, Dongqin Liu, Songlin Hu 0001 |
NLPCC (1) | 3 |
| 2024 | Quartet: A Holistic Hybrid Parallel Framework for Training Large Language Models
Weigang Zhang, Biyu Zhou, Xing Wu 0002, Chaochen Gao, Xuehai Tang, Ruixuan Li 0001, Jizhong Han, Songlin Hu 0001 |
Euro-Par (2) | 3 |
| 2024 | Dial-MAE: ConTextual Masked Auto-Encoder for Retrieval-based Dialogue SystemsabstractZhenpeng Su, Xing W, Wei Zhou, Guangyuan Ma, Songlin Hu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhenpeng Su, Xing Wu 0002, Wei Zhou 0019, Guangyuan Ma, Songlin Hu 0001 |
NAACL-HLT | 2 |
| 2024 | Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage RetrievalabstractMasked auto-encoder pre-training has emerged as a prevalent technique for initializing and enhancing dense retrieval systems. It generally utilizes additional Transformer decoder blocks to provide sustainable supervision signals and compress contextual information into dense representations. However, the underlying reasons for the effectiveness of such a pre-training technique remain unclear. The usage of additional Transformer-based decoders also incurs significant computational costs. In this study, we aim to shed light on this issue by revealing that masked auto-encoder (MAE) pre-training with enhanced decoding significantly improves the term coverage of input tokens in dense representations, compared to vanilla BERT checkpoints. Building upon this observation, we propose a modification to the traditional MAE by replacing the decoder of a masked auto-encoder with a completely simplified Bag-of-Word prediction task. This modification enables the efficient compression of lexical signals into dense representations through unsupervised pre-training. Remarkably, our proposed method achieves state-of-the-art retrieval performance on several large-scale retrieval benchmarks without requiring any additional parameters, which provides a 67% training speed-up compared to standard masked auto-encoder pre-training with enhanced decoding. Guangyuan Ma, Xing Wu 0002, Zijia Lin, Songlin Hu 0001 |
SIGIR | 2 |
| 2023 | ConTextual Masked Auto-Encoder for Dense Passage RetrievalabstractDense passage retrieval aims to retrieve the relevant passages of a query from a large corpus based on dense representations (i.e., vectors) of the query and the passages. Recent studies have explored improving pre-trained language models to boost dense retrieval performance. This paper proposes CoT-MAE (ConTextual Masked Auto-Encoder), a simple yet effective generative pre-training method for dense passage retrieval. CoT-MAE employs an asymmetric encoder-decoder architecture that learns to compress the sentence semantics into a dense vector through self-supervised and context-supervised masked auto-encoding. Precisely, self-supervised masked auto-encoding learns to model the semantics of the tokens inside a text span, and context-supervised masked auto-encoding learns to model the semantical correlation between the text spans. We conduct experiments on large-scale passage retrieval benchmarks and show considerable improvements over strong baselines, demonstrating the high efficiency of CoT-MAE. Our code is available at https://github.com/caskcsg/ir/tree/main/cotmae. Xing Wu 0002, Guangyuan Ma, Zijia Lin, Zhongyuan Wang 0006, Songlin Hu 0001 |
AAAI | 1 |
| 2023 | Query-as-context Pre-training for Dense Passage RetrievalabstractRecently, methods have been developed to improve the performance of dense passage retrieval by using context-supervised pre-training.These methods simply consider two passages from the same document to be relevant, without taking into account the potential negative impacts of weakly correlated pairs.Thus, this paper proposes query-as-context pre-training, a simple yet effective pre-training technique to alleviate the issue.Query-as-context pretraining assumes that the query derived from a passage is more likely to be relevant to that passage and forms a passage-query pair.These passage-query pairs are then used in contrastive or generative context-supervised pre-training.The pre-trained models are evaluated on largescale passage retrieval benchmarks and out-ofdomain zero-shot benchmarks.Experimental results show that query-as-context pre-training brings considerable gains for retrieval performances, demonstrating its effectiveness and efficiency. Xing Wu 0002, Guangyuan Ma, Wanhui Qian, Zijia Lin, Songlin Hu 0001 |
EMNLP | 1 |
| 2022 | ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence EmbeddingabstractContrastive learning has been attracting much attention for learning unsupervised sentence embeddings. The current state-of-the-art unsupervised method is the unsupervised SimCSE (unsup-SimCSE). Unsup-SimCSE takes dropout as a minimal data augmentation method, and passes the same input sentence to a pre-trained Transformer encoder (with dropout turned on) twice to obtain the two corresponding embeddings to build a positive pair. As the length information of a sentence will generally be encoded into the sentence embeddings due to the usage of position embedding in Transformer, each positive pair in unsup-SimCSE actually contains the same length information. And thus unsup-SimCSE trained with these positive pairs is probably biased, which would tend to consider that sentences of the same or similar length are more similar in semantics. Through statistical observations, we find that unsup-SimCSE does have such a problem. To alleviate it, we apply a simple repetition operation to modify the input sentence, and then pass the input sentence and its modified counterpart to the pre-trained Transformer encoder, respectively, to get the positive pair. Additionally, we draw inspiration from the community of computer vision and introduce a momentum contrast, enlarging the number of negative pairs without additional calculations. The proposed two modifications are applied on positive and negative pairs separately, and build a new sentence embedding method, termed Enhanced Unsup-SimCSE (ESimCSE). We evaluate the proposed ESimCSE on several benchmark datasets w.r.t the semantic text similarity (STS) task. Experimental results show that ESimCSE outperforms the state-of-the-art unsup-SimCSE by an average Spearman correlation of 2.02% on BERT-base. Xing Wu 0002, Chaochen Gao, Liangjun Zang, Jizhong Han, Zhongyuan Wang 0006, Songlin Hu 0001 |
COLING | 1 |
| 2022 | Smoothed Contrastive Learning for Unsupervised Sentence EmbeddingabstractUnsupervised contrastive sentence embedding models, e.g., unsupervised SimCSE, use the InfoNCE loss function in training. Theoretically, we expect to use larger batches to get more adequate comparisons among samples and avoid overfitting. However, increasing batch size leads to performance degradation when it exceeds a threshold, which is probably due to the introduction of false-negative pairs through statistical observation. To alleviate this problem, we introduce a simple smoothing strategy upon the InfoNCE loss function, termed Gaussian Smoothed InfoNCE (GS-InfoNCE). In other words, we add random Gaussian noise as an extension to the negative pairs without increasing the batch size. Through experiments on the semantic text similarity tasks, though simple, the proposed smoothing strategy brings improvements to unsupervised SimCSE. Xing Wu 0002, Chaochen Gao, Yipeng Su, Jizhong Han, Zhongyuan Wang 0006, Songlin Hu 0001 |
COLING | 1 |
| 2019 | Imbalanced Sentiment Classification Enhanced with Discourse Marker
Tao Zhang 0101, Xing Wu 0002, Jizhong Han, Songlin Hu 0001 |
ICANN (4) | 2 |
| 2019 | Mask and Infill: Applying Masked Language Model for Sentiment TransferabstractThis paper focuses on the task of sentiment transfer on non-parallel text, which modifies sentiment attributes (e.g., positive or negative) of sentences while preserving their attribute-independent contents. Existing methods adopt RNN encoder-decoder structure to generate a new sentence of a target sentiment word by word, which is trained on a particular dataset from scratch and have limited ability to produce satisfactory sentences. When people convert the sentiment attribute of a given sentence, a simple but effective approach is to only replace the sentiment tokens of the sentence with other expressions indicative of the target sentiment, instead of building a new sentence from scratch. Such a process is very similar to the task of Text Infilling or Cloze. With this intuition, we propose a two steps approach: Mask and Infill. In the \emph{mask} step, we identify and mask the sentiment tokens of a given sentence. In the \emph{infill} step, we utilize a pre-trained Masked Language Model (MLM) to infill the masked positions by predicting words or phrases conditioned on the context\footnote{In this paper, \emph{content} and \emph{context} are equivalent, \emph{style}, \emph{attribute} and \emph{label} are equivalent.}and target sentiment. We evaluate our model on two review datasets \emph{Yelp} and \emph{Amazon} by quantitative, qualitative, and human evaluations. Experimental results demonstrate that our model achieve state-of-the-art performance on both accuracy and BLEU scores. Xing Wu 0002, Tao Zhang 0101, Liangjun Zang, Jizhong Han, Songlin Hu 0001 |
IJCAI | 1 |