VLDB 2026 Research / reviewers in the wild / expert
Yunshui Li
dblp:339/2740
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hierarchical Context Pruning: Optimizing Real-World Code Completion with Repository-Level Pretrained Code LLMsabstractSome of the latest released Code Large Language Models (Code LLMs) have been trained on repository-level code data, enabling them to perceive repository structures and utilize cross-file code information. This capability allows us to directly concatenate the content of repository code files in prompts to achieve repository-level code completion. However, in real development scenarios, directly concatenating all code repository files in a prompt can easily exceed the context window of Code LLMs, leading to a significant decline in completion performance. Additionally, overly long prompts can increase completion latency, negatively impacting the user experience. In this study, we conducted extensive experiments, including completion error analysis, topology dependency analysis, and cross-file content analysis, to investigate the factors affecting repository-level code completion. Based on the conclusions drawn from these preliminary experiments, we proposed a strategy called **Hierarchical Context Pruning (HCP)** to construct high-quality completion prompts. We applied the **HCP** to six Code LLMs and evaluated them on the CrossCodeEval dataset. The experimental results showed that, compared to previous methods, the prompts constructed using our **HCP** strategy achieved higher completion accuracy on five out of six Code LLMs. Additionally, the **HCP** managed to keep the prompt length around 8k tokens (whereas the full repository code is approximately 50k tokens), significantly improving completion throughput. Our code and data will be publicly available. Lei Zhang 0201, Yunshui Li, Jiaming Li 0004, Xiaobo Xia, Jiaxi Yang 0004, Run Luo, Minzheng Wang 0001, Longze Chen, Junhao Liu 0001, Qiang Qu 0001, Min Yang 0007 |
AAAI | 2 |
| 2025 | GATEAU: Selecting Influential Samples for Long Context AlignmentabstractShuzheng Si, Haozhe Zhao, Gang Chen, Yunshui Li, Kangyang Luo, Chuancheng Lv, Kaikai An, Fanchao Qi, Baobao Chang, Maosong Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Shuzheng Si, Haozhe Zhao, Gang Chen 0039, Yunshui Li, Kangyang Luo, Chuancheng Lv, Kaikai An, Fanchao Qi, Baobao Chang, Maosong Sun 0001 |
EMNLP | 4 |
| 2025 | DEEM: Diffusion models serve as the eyes of large language models for image perceptionabstractThe development of large language models (LLMs) has significantly advanced the emergence of large multimodal models (LMMs). While LMMs have achieved tremendous success by promoting the synergy between multimodal comprehension and creation, they often face challenges when confronted with out-of-distribution data, such as which can hardly distinguish orientation, quantity, color, structure, etc. This is primarily due to their reliance on image encoders trained to encode images into task-relevant features, which may lead them to disregard irrelevant details. Delving into the modeling capabilities of diffusion models for images naturally prompts the question: Can diffusion models serve as the eyes of large language models for image perception? In this paper, we propose DEEM, a simple but effective approach that utilizes the generative feedback of diffusion models to align the semantic distributions of the image encoder. This addresses the drawbacks of previous methods that solely relied on image encoders like CLIP-ViT, thereby enhancing the model's resilience against out-of-distribution samples and reducing visual hallucinations. Importantly, this is achieved without requiring additional training modules and with fewer training parameters. We extensively evaluated DEEM on both our newly constructed RobustVQA benchmark and other well-known benchmarks, POPE and MMVP, for visual hallucination and perception. In particular, DEEM improves LMM's visual perception performance to a large extent (e.g., 4\% ↑ on RobustVQA, 6.5\% ↑ on MMVP and 12.8 \% ↑ on POPE ). Compared to the state-of-the-art interleaved content generation models, DEEM exhibits enhanced robustness and a superior capacity to alleviate model hallucinations while utilizing fewer trainable parameters, less pre-training data (10\%), and a smaller base model size. Extensive experiments demonstrate that DEEM enhances the performance of LMMs on various downstream tasks without inferior performance in the long term, including visual question answering, image captioning, and text-conditioned image synthesis. Run Luo, Yunshui Li, Longze Chen, Wanwei He, Ting-En Lin, Lei Zhang 0201, Zikai Song, Hamid Alinejad-Rokny, Xiaobo Xia, Tongliang Liu, Binyuan Hui, Min Yang 0007 |
ICLR | 2 |
| 2025 | Model Merging in Pre-training of Large Language ModelsabstractModel merging has emerged as a promising technique for enhancing large language models, though its application in large-scale pre-training remains relatively unexplored. In this paper, we present a comprehensive investigation of model merging techniques during the pre-training process. Through extensive experiments with both dense and Mixture-of-Experts (MoE) architectures ranging from millions to over 100 billion parameters, we demonstrate that merging checkpoints trained with constant learning rates not only achieves significant performance improvements but also enables accurate prediction of annealing behavior. These improvements lead to both more efficient model development and significantly lower training costs. Our detailed ablation studies on merging strategies and hyperparameters provide new insights into the underlying mechanisms while uncovering novel applications. Through comprehensive experimental analysis, we offer the open-source community practical pre-training guidelines for effective model merging. Yunshui Li, Yiyuan Ma, Chaoyi Zhang, Jianqiao Lu, Ziwen Xu, Mengzhao Chen, Minrui Wang, Shiyi Zhan, Xunhao Lai, Yao Luo, Xingyan Bin, Hongbin Ren, Mingji Han, Wenhao Hao, Bairen Yi, LingJun Liu, Bole Ma, Xiaoying Jia 0005 |
NeurIPS | 1 |
| 2024 | Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language ModelsabstractLongze Chen, Ziqiang Liu, Wanwei He, Yinhe Zheng, Hao Sun, Yunshui Li, Run Luo, Min Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Longze Chen, Wanwei He, Yinhe Zheng, Yunshui Li, Run Luo, Min Yang 0007 |
ACL (1) | 6 |
| 2024 | One-Shot Learning as Instruction Data Prospector for Large Language ModelsabstractYunshui Li, Binyuan Hui, Xiaobo Xia, Jiaxi Yang, Min Yang, Lei Zhang, Shuzheng Si, Ling-Hao Chen, Junhao Liu, Tongliang Liu, Fei Huang, Yongbin Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yunshui Li, Binyuan Hui, Xiaobo Xia, Jiaxi Yang 0004, Min Yang 0007, Lei Zhang 0201, Shuzheng Si, Junhao Liu 0001, Tongliang Liu, Fei Huang 0002, Yongbin Li 0001 |
ACL (1) | 1 |
| 2024 | Marathon: A Race Through the Realm of Long Context with Large Language ModelsabstractLei Zhang, Yunshui Li, Ziqiang Liu, Jiaxi Yang, Junhao Liu, Longze Chen, Run Luo, Min Yang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Lei Zhang 0201, Yunshui Li, Jiaxi Yang 0004, Junhao Liu 0001, Longze Chen, Run Luo, Min Yang 0007 |
ACL (1) | 2 |
| 2024 | TP-Link: Fine-grained Pre-Training for Text-to-SQL Parsing with Linking InformationabstractIn this paper, we introduce an innovative pre-training framework TP-Link, which aims to improve context-dependent Text-to-SQL Parsing by leveraging Linking information. This enhancement is achieved through better representation of both natural language utterances and the database schema, ultimately facilitating more effective text-to-SQL conversations. We present two novel pre-training objectives: (i) utterance linking prediction (ULP) task that models intricate syntactic relationships among natural language utterances in context-dependent text-to-SQL scenarios, and (ii) schema linking prediction (SLP) task that focuses on capturing fine-grained schema linking relationships between the utterances and the database schema. Extensive experiments demonstrate that our proposed TP-Link achieves state-of-the-art performance on two leading downstream benchmarks (i.e., SParC and CoSQL). Shujie Li 0001, Zefeng Cai, Yunshui Li, Chengming Li 0004, Xiping Hu, Ruifeng Xu 0001, Min Yang 0007 |
LREC/COLING | 5 |
| 2024 | Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QAabstractMinzheng Wang, Longze Chen, Fu Cheng, Shengyi Liao, Xinghua Zhang, Bingli Wu, Haiyang Yu, Nan Xu, Lei Zhang, Run Luo, Yunshui Li, Min Yang, Fei Huang, Yongbin Li. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Minzheng Wang 0001, Longze Chen, Fu Cheng, Shengyi Liao, Xinghua Zhang 0001, Bingli Wu, Haiyang Yu 0003, Nan Xu 0004, Lei Zhang 0201, Run Luo, Yunshui Li, Min Yang 0007, Fei Huang 0002, Yongbin Li 0001 |
EMNLP | 11 |
| 2023 | PaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional ExpertsabstractPerceiving multi-modal information and fulfilling dialogues with humans is a long-term goal of artificial intelligence.Pre-training is commonly regarded as an effective approach for multi-modal dialogue.However, due to the limited availability of multi-modal dialogue data, there is still scarce research on multi-modal dialogue pre-training.Yet another intriguing challenge emerges from the encompassing nature of multi-modal dialogue, which involves various modalities and tasks.Moreover, new forms of tasks may arise at unpredictable points in the future.Hence, it is essential for designed multi-modal dialogue models to possess sufficient flexibility to adapt to such scenarios.This paper proposes PaCE, a unified, structured, compositional multi-modal dialogue pretraining framework.It utilizes a combination of several fundamental experts to accommodate multiple dialogue-related tasks and can be pre-trained using limited dialogue and extensive non-dialogue multi-modal data.Furthermore, we propose a progressive training method where old experts from the past can assist new experts, facilitating the expansion of their capabilities.Experimental results demonstrate that PaCE achieves state-of-the-art results on eight multi-modal dialog benchmarks. Yunshui Li, Binyuan Hui, Zhichao Yin, Min Yang 0007, Fei Huang 0002, Yongbin Li 0001 |
ACL (1) | 1 |