VLDB 2026 Research / reviewers in the wild / expert
Daixuan Cheng
dblp:289/2865
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-0405-9707ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reasoning with Exploration: An Entropy PerspectiveabstractBalancing exploration and exploitation is a central goal in reinforcement learning (RL). Despite recent advances in enhancing language model (LM) reasoning, most methods lean toward exploitation, and increasingly encounter performance plateaus. In this work, we revisit entropy -- a signal of exploration in RL -- and examine its relationship to exploratory reasoning in LMs. Through empirical analysis, we uncover positive correlations between high-entropy regions and three types of exploratory reasoning actions: (1) pivotal tokens that determine or connect logical steps, (2) reflective actions such as self-verification and correction, and (3) rare behaviors under-explored by the base LMs. Motivated by this, we introduce a minimal modification to standard RL with only one line of code: augmenting the advantage function with an entropy-based term. Unlike traditional maximum-entropy methods which encourage exploration by promoting uncertainty, we encourage exploration by promoting deeper and longer reasoning chains. Notably, our method achieves significant gains on the Pass@K metric -- an upper-bound estimator of LM reasoning capabilities -- even when evaluated with extremely large K values, pushing the boundaries of LM reasoning. Daixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai 0026, Wayne Xin Zhao, Furu Wei |
AAAI | 1 |
| 2025 | How to Synthesize Text Data without Model Collapse?abstractModel collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance. With the proliferation of AI models, synthetic data will fundamentally reshape the web data ecosystem. Future GPT-$\{n\}$ models will inevitably be trained on a blend of synthetic and human-produced data. In this paper, we focus on two questions: what is the impact of synthetic data on language model training, and how to synthesize data without model collapse? We first pre-train language models across different proportions of synthetic data, revealing a negative correlation between the proportion of synthetic data and model performance. We further conduct statistical analysis on synthetic data to uncover distributional shift phenomenon and over-concentration of n-gram features. Inspired by the above findings, we propose token editing on human-produced data to obtain semi-synthetic data. As a proof of concept, we theoretically demonstrate that token-level editing can prevent model collapse, as the test error is constrained by a finite upper bound. We conduct extensive experiments on pre-training from scratch, continual pre-training, and supervised fine-tuning. The results validate our theoretical proof that token-level editing improves data quality and enhances model performance. Xuekai Zhu, Daixuan Cheng, Hengli Li, Ermo Hua, Xingtai Lv, Ning Ding 0002, Zhouhan Lin, Zilong Zheng, Bowen Zhou 0002 |
ICML | 2 |
| 2025 | Federated Fine-Tuning on Heterogeneous LoRAs With Error-Compensated AggregationabstractFederated learning (FL) has recently been applied to the parameter-efficient fine-tuning (PEFT) of large language models (LLMs). While promising, client resource heterogeneity has imposed the challenge of the "bucket effect" to FL, where model configuration must cater to the client with the fewest resources. To tackle this issue, heterogeneous low-rank adaptation (LoRA) has recently emerged in FL, which enables clients to do local fine-tuning with different LoRA ranks. However, existing works in this area typically adopt zero-padding, stacking, or singular value decomposition (SVD) for LoRA aggregation, which often incur precision loss or significant overhead, limiting their practicality. In this article, we propose ECLoRA, a novel method for federated fine-tuning with heterogeneous LoRA settings across clients. ECLoRA employs randomized SVD (RSVD) to dramatically reduce aggregation overhead while introducing an error compensation (EC) mechanism that incorporates the decomposition error from previous rounds to improve aggregation precision. Extensive experiments on four widely used foundation models across six public tasks demonstrate the effectiveness of ECLoRA. Specifically, ECLoRA is: (1) accurate, significantly improving the final model performance; (2) fast, accelerating convergence with an average speedup of $1.54\times $ to $3.01\times $ ; and (3) practical, reducing aggregation time by approximately $40\times $ compared to classical SVD. Wanyi Ning, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Daixuan Cheng, Cong Liu 0046, Lei Zhang 0094, Zirui Zhuang, Jianxin Liao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Instruction Pre-Training: Language Models are Supervised Multitask LearnersabstractUnsupervised multitask pre-training has been the critical method behind the recent success of language models (LMs).However, supervised multitask learning still holds significant promise, as scaling it in the post-training stage trends towards better generalization.In this paper, we explore supervised multitask pretraining by proposing Instruction Pre-Training, a framework that scalably augments massive raw corpora with instruction-response pairs to pre-train LMs.The instruction-response pairs are generated by an efficient instruction synthesizer built on open-source models.In our experiments, we synthesize 200M instruction-response pairs covering 40+ task categories to verify the effectiveness of Instruction Pre-Training.In pre-training from scratch, Instruction Pre-Training not only consistently enhances pre-trained base models but also benefits more from further instruction tuning.In continual pre-training, Instruction Pre-Training enables Llama3-8B to be comparable to or even outperform Llama3-70B.Our model, code, and data are available at https://github.com/microsoft/LMOps.Ins: When is the finale of season 7? Let's think step by step. Daixuan Cheng, Yuxian Gu, Shaohan Huang, Junyu Bi, Minlie Huang, Furu Wei |
EMNLP | 1 |
| 2024 | Adapting Large Language Models via Reading ComprehensionabstractWe explore how continued pre-training on domain-specific corpora influences large language models, revealing that training on the raw corpora endows the model with domain knowledge, but drastically hurts its prompting ability for question answering. Taken inspiration from human learning via reading comprehension--practice after reading improves the ability to answer questions based on the learned knowledge--we propose a simple method for transforming raw corpora into reading comprehension texts. Each raw text is enriched with a series of tasks related to its content. Our method, highly scalable and applicable to any pre-training corpora, consistently enhances performance across various tasks in three different domains: biomedicine, finance, and law. Notably, our 7B language model achieves competitive performance with domain-specific models of much larger scales, such as BloombergGPT-50B. Furthermore, we demonstrate that domain-specific reading comprehension texts can improve the model's performance even on general benchmarks, showing the potential to develop a general model across even more domains. Our model, code, and data are available at https://github.com/microsoft/LMOps. Daixuan Cheng, Shaohan Huang, Furu Wei |
ICLR | 1 |
| 2024 | MDR: Model-Specific Demonstration Retrieval at Inference Time for In-Context LearningabstractHuazheng Wang, Jinming Wu, Haifeng Sun, Zixuan Xia, Daixuan Cheng, Jingyu Wang, Qi Qi, Jianxin Liao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Huazheng Wang, Haifeng Sun 0001, Zixuan Xia, Daixuan Cheng, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
NAACL-HLT | 5 |
| 2024 | FM-Delta: Lossless Compression for Storing Massive Fine-tuned Foundation ModelsabstractPre-trained foundation models, particularly large language models, have achieved remarkable success and led to massive fine-tuned variants. These models are commonly fine-tuned locally and then uploaded by users to cloud platforms such as HuggingFace for secure storage. However, the huge model number and their billion-level parameters impose heavy storage overhead for cloud with limited resources. Our empirical and theoretical analysis reveals that most fine-tuned models in cloud have a small difference (delta) from their pre-trained models. To this end, we propose a novel lossless compression scheme FM-Delta specifically for storing massive fine-tuned models in cloud. FM-Delta maps fine-tuned and pre-trained model parameters into integers with the same bits, and entropy codes their integer delta. In this way, cloud only needs to store one uncompressed pre-trained model and other compressed fine-tuned models.
Extensive experiments have demonstrated that FM-Delta efficiently reduces cloud storage consumption for massive fine-tuned models by an average of around 50% with only negligible additional time in most end-to-end cases. For example, on up to 10 fine-tuned models in the GPT-NeoX-20B family, FM-Delta reduces the original storage requirement from 423GB to 205GB, significantly saving cloud storage costs. Wanyi Ning, Jingyu Wang 0001, Qi Qi 0001, Mengde Zhu, Haifeng Sun 0001, Daixuan Cheng, Jianxin Liao |
NeurIPS | 6 |
| 2023 | How Does Diffusion Influence Pretrained Language Models on Out-of-Distribution Data?abstractTransformer-based pretrained language models (PLMs) have achieved great success in modern NLP. An important advantage of PLMs is good out-of-distribution (OOD) robustness. Recently, diffusion models have attracted a lot of work to apply diffusion to PLMs. It remains under-explored how diffusion influences PLMs on OOD data. The core of diffusion models is a forward diffusion process which gradually applies Gaussian noise to inputs, and a reverse denoising process which removes noise. The noised input reconstruction is a fundamental ability of diffusion models. We directly analyze OOD robustness by measuring the reconstruction loss, including testing the abilities to reconstruct OOD data, and to detect OOD samples. Experiments are conducted by analyzing different training parameters and data statistical features on eight datasets. It shows that finetuning PLMs with diffusion degrades the reconstruction ability on OOD data. The comparison also shows that diffusion models can effectively detect OOD samples, achieving state-of-the-art performance in most of the datasets with an absolute accuracy improvement up to 18%. These results indicate that diffusion reduces OOD robustness of PLMs. Huazheng Wang, Daixuan Cheng, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao, Jing Wang 0039, Cong Liu 0046 |
ECAI | 2 |
| 2023 | UPRISE: Universal Prompt Retrieval for Improving Zero-Shot EvaluationabstractDaixuan Cheng, Shaohan Huang, Junyu Bi, Yuefeng Zhan, Jianfeng Liu, Yujing Wang, Hao Sun, Furu Wei, Weiwei Deng, Qi Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Daixuan Cheng, Shaohan Huang, Junyu Bi, Yuefeng Zhan, Yujing Wang 0002, Hao Sun 0015, Furu Wei, Qi Zhang 0066 |
EMNLP | 1 |
| 2023 | VL-Match: Enhancing Vision-Language Pretraining with Token-Level and Instance-Level MatchingabstractVision-Language Pretraining (VLP) has significantly improved the performance of various vision-language tasks with the matching of images and texts. In this paper, we propose VL-Match, a Vision-Language framework with Enhanced Token-level and Instance-level Matching. At the token level, a Vision-Language Replaced Token Detection task is designed to boost the substantial interaction between text tokens and images, where the text encoder of VLP works as a generator to generate a corrupted text, and the multimodal encoder of VLP works as a discriminator to predict whether each text token in the corrupted text matches the image. At the instance level, in the Image-Text Matching task that judges whether an image-text pair is matched, we propose a novel bootstrapping method to generate hard negative text samples that are different from the positive ones only at the token level. In this way, we can force the network to detect fine-grained differences between images and texts. Notably, with a smaller amount of parameters, VL-Match significantly outperforms previous SOTA on all image-text retrieval tasks. Junyu Bi, Daixuan Cheng, Ping Yao, Bochen Pang, Yuefeng Zhan, Chuanguang Yang, Yujing Wang 0002, Hao Sun 0015, Qi Zhang 0066 |
ICCV | 2 |
| 2021 | Parallel Decoders Guided Lexically Constrained Response GenerationabstractResponse generation is a fundamental function in conversational systems, where controllability of the response is a key problem. In this paper, we consider how to control the response by lexical constraints, namely lexically constrained response generation. The stochastic search-based methods have achieved promising performance in satisfying lexical constraints. The idea of these methods is modifying a sentence through the actions of insertion, deletion and replacement guided by an optimization algorithm. The core of our method is modifying the response by incorporating the lexical constraints and preserving the message-related parts. For this purpose, we propose the novel Parallel Decoders to guide response modification. The first decoder generates responses according to the given message and constraints. The second decoder calculates the relevance between the input and the response. Based on Parallel Decoders, during the modification, we could sample the positions in the response for editing according to the relevance score. Experiments show the proposed framework achieves better performance than the state-of-the-art generation models in terms of constraint relevance, sentence fluency, response diversity and human evaluation. Daixuan Cheng, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001 |
IEEE BigData | 1 |
| 2021 | Spatial-aware stacked regression network for real-time 3D hand pose estimation
Pengfei Ren 0001, Haifeng Sun 0001, Weiting Huang, Jiachang Hao, Daixuan Cheng, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
Neurocomputing | 5 |
| 2021 | Pattern and content controlled response generation
Haifeng Sun 0001, Daixuan Cheng, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
Inf. Process. Manag. | 2 |