VLDB 2026 Research / reviewers in the wild / expert
Zican Dong
dblp:336/7105
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 56% Efficient and distributed learning · 15% Question answering and dialogue systems · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Theoretical computer science
1 paper |
Automated reasoning and model checking · 100% |
Topics — the 22 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model |
2.8 | 4 | 2025 | Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations · NeurIPS 2025 Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework · EMNLP 2025 Exploring Context Window of Large Language Models via Decomposed Positional Vectors · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › data-efficient learning
data-efficient pretraining |
0.9 | 1 | 2025 | YuLan-Mini: Pushing the Limits of Open Data-efficient Language Model · ACL (1) 2025 |
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
expert pruning |
0.9 | 1 | 2025 | Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations · NeurIPS 2025 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model |
0.9 | 1 | 2025 | LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation · ACL (1) 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations · NeurIPS 2025 |
Natural language and speech › Question answering and dialogue systems › knowledge-intensive question answering
multi-document question answering |
0.9 | 1 | 2025 | CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA Capability · EMNLP 2025 |
Machine learning › Representation and self-supervised learning
pre-training |
0.9 | 1 | 2025 | YuLan-Mini: Pushing the Limits of Open Data-efficient Language Model · ACL (1) 2025 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.9 | 1 | 2025 | CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA Capability · EMNLP 2025 |
Natural language and speech › Language models and text generation
test-time scaling |
0.9 | 1 | 2025 | Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework · EMNLP 2025 |
Information retrieval › efficient retrieval
coarse-to-fine retrieval |
0.9 | 1 | 2025 | CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA Capability · EMNLP 2025 |
Automated reasoning and model checking › automated reasoning
mathematical reasoning |
0.9 | 1 | 2025 | Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework · EMNLP 2025 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling
context window extension |
0.8 | 1 | 2024 | Exploring Context Window of Large Language Models via Decomposed Positional Vectors · NeurIPS 2024 |
Natural language and speech › Language models and text generation › compositional generalization › length generalization
length extrapolation |
0.8 | 1 | 2024 | Exploring Context Window of Large Language Models via Decomposed Positional Vectors · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
positional encoding |
0.8 | 1 | 2024 | Exploring Context Window of Large Language Models via Decomposed Positional Vectors · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Exploring Context Window of Large Language Models via Decomposed Positional Vectors · NeurIPS 2024 |
Natural language and speech › Question answering and dialogue systems
knowledge base question answering |
0.7 | 1 | 2023 | StructGPT: A General Framework for Large Language Model to Reason over Structured Data · EMNLP 2023 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.7 | 1 | 2023 | StructGPT: A General Framework for Large Language Model to Reason over Structured Data · EMNLP 2023 |
Natural language and speech › Question answering and dialogue systems
question answering over structured data |
0.7 | 1 | 2023 | StructGPT: A General Framework for Large Language Model to Reason over Structured Data · EMNLP 2023 |
Natural language and speech › Language models and text generation › agentic language model › tool-augmented language models
tool-augmented reasoning |
0.7 | 1 | 2023 | StructGPT: A General Framework for Large Language Model to Reason over Structured Data · EMNLP 2023 |
Natural language and speech › Language models and text generation › model steering › language model steering
attention steering |
0.3 | 1 | 2025 | CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA Capability · EMNLP 2025 |
Machine learning › Reinforcement learning
imitation learning |
0.3 | 1 | 2025 | Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework · EMNLP 2025 |
Natural language and speech › Language models and text generation
self-improvement |
0.3 | 1 | 2025 | Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling Framework · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
self-consistency · 1.7retrieval-augmented generation · 1.7reinforcement learning · 1.7imitation learning · 1.7attention steering · 1.7retrieval head · 0.9restoration distillation · 0.9knowledge distillation · 0.9data synthesis · 0.9data selection · 0.9data curriculum · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Survey of Large Language ModelsabstractAbstract The rapid evolution of large language models (LLMs) has driven a transformative shift in artificial intelligence (AI), reshaping both research paradigms and practical applications. Distinguished from their predecessors by unprecedented scale and advanced capabilities, LLMs necessitate new frameworks for understanding their development, behavior, and societal impact. This survey systematically reviews recent advancements in LLM techniques across four key dimensions: (1) pre-training methodologies, which establish core model capabilities through large-scale self-supervised training, architectural innovations, and data curation strategies; (2) post-training techniques, including supervised fine-tuning and reinforcement learning, which adapt foundational models to downstream tasks and enhance their alignment and safety; (3) utilization strategies, such as in-context learning, prompt engineering, and agentic reasoning, that optimize real-world deployment and enable effective interaction with external environments; and (4) evaluation methods, encompassing benchmarks for key ability dimensions such as core language capabilities, reasoning, and safety, which support comprehensive and reliable assessment of model performance. Additionally, we identify critical research issues, including those concerning theoretical foundations, efficient scaling, alignment, and agentic capability, and highlight the open challenges they present. By synthesizing state-of-the-art insights and emerging trends, this survey aims to provide a systematic and comprehensive framework for understanding the trajectory, current limitations, and future directions of LLM progress. Wayne Xin Zhao, Kun Zhou 0002, Junyi Li 0001, Zican Dong, Yupeng Hou, Beichen Zhang 0003, Yingqian Min, Junjie Zhang 0009, Peiyu Liu 0002, Xiaolei Wang 0005, Yifan Du 0002, Chen Yang 0032, Zhipeng Chen 0001, Jinhao Jiang, Ruiyang Ren, Yifan Li 0009, Xinyu Tang 0004, Zikang Liu 0001, Jian-Yun Nie, Ji-Rong Wen |
Frontiers Comput. Sci. | 5 |
| 2025 | LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration DistillationabstractZican Dong, Junyi Li, Jinhao Jiang, Mingyu Xu, Xin Zhao, Bingning Wang, Weipeng Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zican Dong, Jinhao Jiang, Bingning Wang, Weipeng Chen |
ACL (1) | 1 |
| 2025 | YuLan-Mini: Pushing the Limits of Open Data-efficient Language ModelabstractDue to the immense resource demands and the involved complex techniques, it is still challenging for successfully pre-training a large language models (LLMs) with state-of-the-art performance. In this paper, we explore the key bottlenecks and designs during pre-training, and make the following contributions: (1) a comprehensive investigation into the factors contributing to training instability; (2) a robust optimization approach designed to mitigate training instability effectively; (3) an elaborate data pipeline that integrates data synthesis, data curriculum, and data selection. By integrating the above techniques, we create a rather low-cost training recipe and use it to pre-train YuLan-Mini, a fully-open base model with 2.4B parameters on 1.08T tokens. Remarkably, YuLan-Mini achieves top-tier performance among models of similar parameter scale, with comparable performance to industry-leading models that require significantly more data. To facilitate reproduction, we release the full details of training recipe and data composition. Project details can be accessed at the following link: https://anonymous.4open.science/r/YuLan-Mini/README.md. Huatong Song, Jie Chen 0007, Kun Zhou 0002, Yutao Zhu 0001, Jinhao Jiang, Zican Dong, Xu Miao, Wayne Xin Zhao, Ji-Rong Wen |
ACL (1) | 9 |
| 2025 | Sticker-TTS: Learn to Utilize Historical Experience with a Sticker-driven Test-Time Scaling FrameworkabstractLarge reasoning models (LRMs) have exhibited strong performance on complex reasoning tasks, with further gains achievable through increased computational budgets at inference.However, current test-time scaling methods predominantly rely on redundant sampling, ignoring the historical experience utilization, thereby limiting computational efficiency.To overcome this limitation, we propose Sticker-TTS, a novel test-time scaling framework that coordinates three collaborative LRMs to iteratively explore and refine solutions guided by historical attempts.At the core of our framework are distilled key conditions-termed stickers-which drive the extraction, refinement, and reuse of critical information across multiple rounds of reasoning.To further enhance the efficiency and performance of our framework, we introduce a two-stage optimization strategy that combines imitation learning with self-improvement, enabling progressive refinement.Extensive evaluations on three challenging mathematical reasoning benchmarks, including AIME-24, AIME-25, and Olym-MATH, demonstrate that Sticker-TTS consistently surpasses strong baselines, including self-consistency and advanced reinforcement learning approaches, under comparable inference budgets.These results highlight the effectiveness of sticker-guided historical experience utilization. Jie Chen 0007, Jinhao Jiang, Yingqian Min, Zican Dong, Wayne Xin Zhao, Ji-Rong Wen |
EMNLP | 4 |
| 2025 | CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA CapabilityabstractAdvancements in Large Language Models (LLMs) have extended their input context length, yet they still struggle with retrieval and reasoning in long-context inputs.Existing methods propose to utilize the prompt strategy and Retrieval-Augmented Generation (RAG) to alleviate this limitation.However, they still face challenges in balancing retrieval precision and recall, impacting their efficacy in answering questions.To address this, we introduce CAFE, a two-stage coarse-to-fine method to enhance multi-document questionanswering capacities.By gradually eliminating the negative impacts of background and distracting documents, CAFE makes the responses more reliant on the evidence documents.Initially, a coarse-grained filtering method leverages retrieval heads to identify and rank relevant documents.Then, a fine-grained steering method guides attention to the most relevant content.Experiments across benchmarks show that CAFE outperforms baselines, achieving an average SubEM improvement of up to 22.1% and 13.7% over SFT and RAG methods, respectively, across three different models.Our code is available at https://github. com/RUCAIBox/CAFE. Jinhao Jiang, Zican Dong, Wayne Xin Zhao |
EMNLP | 3 |
| 2025 | Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot DemonstrationsabstractMixture-of-Experts (MoE) models achieve a favorable trade-off between performance and inference efficiency by activating only a subset of experts. However, the memory overhead of storing all experts remains a major limitation, especially in large-scale MoE models such as DeepSeek-R1 (671B).
In this study, we investigate domain specialization and expert redundancy in large-scale MoE models and uncover a consistent behavior we term~\emph{few-shot expert localization}, with only a few in-domain demonstrations, the model consistently activates a sparse and stable subset of experts on tasks within the same domain.
Building on this observation, we propose a simple yet effective pruning framework, \textbf{EASY-EP}, that leverages a few domain-specific demonstrations to identify and retain only the most relevant experts.
EASY-EP comprises two key components: \textbf{output-aware expert importance assessment} and \textbf{expert-level token contribution estimation}. The former evaluates the importance of each expert for the current token by considering the gating scores and L2 norm of the outputs of activated experts, while the latter assesses the contribution of tokens based on representation similarities before and after routed experts.
Experiments on DeepSeek-R1 and DeepSeek-V3-0324 show that our method can achieve comparable performances and $2.99\times$ throughput under the same memory budget as the full model, with only half the experts. Our code is available at https://github.com/RUCAIBox/EASYEP. Zican Dong, Peiyu Liu 0002, Wayne Xin Zhao |
NeurIPS | 1 |
| 2024 | BAMBOO: A Comprehensive Benchmark for Evaluating Long Text Modeling Capacities of Large Language ModelsabstractLarge language models (LLMs) have achieved dramatic proficiency over NLP tasks with normal length. Recently, multiple studies have committed to extending the context length and enhancing the long text modeling capabilities of LLMs. To comprehensively evaluate the long context ability of LLMs, we propose BAMBOO, a multi-task long context benchmark. BAMBOO has been designed with four principles: comprehensive capacity evaluation, avoidance of data contamination, accurate automatic evaluation, and different length levels. It consists of 10 datasets from 5 different long text understanding tasks, i.e., question answering, hallucination detection, text sorting, language modeling, and code completion, to cover various domains and core capacities of LLMs. We conduct experiments with five widely-used long-context models and further discuss five key questions for long text research. In the end, we discuss problems of current long-context models and point out future directions for enhancing long text modeling capacities. We release our data, prompts, and code at https://anonymous.4open.science/r/BAMBOO/. Zican Dong, Junyi Li 0001, Wayne Xin Zhao, Ji-Rong Wen |
LREC/COLING | 1 |
| 2024 | Exploring Context Window of Large Language Models via Decomposed Positional VectorsabstractTransformer-based large language models (LLMs) typically have a limited context window, resulting in significant performance degradation when processing text beyond the length of the context window. Extensive studies have been proposed to extend the context window and achieve length extrapolation of LLMs, but there is still a lack of in-depth interpretation of these approaches. In this study, we explore the positional information within and beyond the context window for deciphering the underlying mechanism of LLMs. By using a mean-based decomposition method, we disentangle positional vectors from hidden states of LLMs and analyze their formation and effect on attention. Furthermore, when texts exceed the context window, we analyze the change of positional vectors in two settings, i.e., direct extrapolation and context window extension. Based on our findings, we design two training-free context window extension methods, positional vector replacement and attention window extension. Experimental results show that our methods can effectively extend the context window length. Zican Dong, Junyi Li 0001, Xin Men, Wayne Xin Zhao, Bingning Wang, Zhen Tian 0001, Weipeng Chen, Ji-Rong Wen |
NeurIPS | 1 |
| 2023 | StructGPT: A General Framework for Large Language Model to Reason over Structured DataabstractIn this paper, we aim to improve the reasoning ability of large language models (LLMs) over structured data in a unified way.Inspired by the studies on tool augmentation for LLMs, we develop an Iterative Reading-then-Reasoning (IRR) framework to solve question answering tasks based on structured data, called StructGPT.In this framework, we construct the specialized interfaces to collect relevant evidence from structured data (i.e., reading), and let LLMs concentrate on the reasoning task based on the collected information (i.e., reasoning).Specially, we propose an invokinglinearization-generation procedure to support LLMs in reasoning on the structured data with the help of the interfaces.By iterating this procedure with provided interfaces, our approach can gradually approach the target answers to a given query.Experiments conducted on three types of structured data show that StructGPT greatly improves the performance of LLMs, under the few-shot and zero-shot settings.Our codes and data are publicly available at https://github.com/RUCAIBox/StructGPT. Jinhao Jiang, Kun Zhou 0002, Zican Dong, Keming Ye, Wayne Xin Zhao, Ji-Rong Wen |
EMNLP | 3 |