VLDB 2026 Research / reviewers in the wild / expert
Kai Lv 0001
dblp:191/2440-1
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 48% Efficient and distributed learning · 22% Planning, search and constraint satisfaction · 6% | |
| Databases, data mining, and information retrieval
2 papers |
Data integration and cleaning · 82% Information retrieval · 18% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
attention mechanism |
0.9 | 1 | 2025 | ReAttention: Training-Free Infinite Context with Finite Attention Scope · ICLR 2025 |
Natural language and speech › Language models and text generation › large language model inference
hybrid inference |
0.9 | 1 | 2025 | Firewall Routing: Blocking Leads to Better Hybrid Inference for LLMs · EMNLP 2025 |
Machine learning › Efficient and distributed learning › inference efficiency
inference optimization |
0.9 | 1 | 2025 | Firewall Routing: Blocking Leads to Better Hybrid Inference for LLMs · EMNLP 2025 |
Natural language and speech › Language models and text generation
large language model inference |
0.9 | 1 | 2025 | Firewall Routing: Blocking Leads to Better Hybrid Inference for LLMs · EMNLP 2025 |
Natural language and speech › Language models and text generation › compositional generalization › length generalization
length extrapolation |
0.9 | 1 | 2025 | ReAttention: Training-Free Infinite Context with Finite Attention Scope · ICLR 2025 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model |
0.9 | 1 | 2025 | ReAttention: Training-Free Infinite Context with Finite Attention Scope · ICLR 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search |
0.9 | 1 | 2025 | FastMCTS: A Simple Sampling Strategy for Data Synthesis · ACL (1) 2025 |
Machine learning › Probabilistic and Bayesian machine learning
sampling |
0.9 | 1 | 2025 | FastMCTS: A Simple Sampling Strategy for Data Synthesis · ACL (1) 2025 |
Data integration and cleaning
data quality |
0.9 | 1 | 2025 | CritiQ: Mining Data Quality Criteria from Human Preferences · ACL (1) 2025 |
Machine learning › Transfer learning and domain adaptation › fine-tuning
full parameter fine-tuning |
0.8 | 1 | 2024 | Full Parameter Fine-tuning for Large Language Models with Limited Resources · ACL (1) 2024 |
Machine learning › Optimization for machine learning › optimization › optimizer design
memory-efficient optimizer |
0.8 | 1 | 2024 | Full Parameter Fine-tuning for Large Language Models with Limited Resources · ACL (1) 2024 |
Machine learning › Efficient and distributed learning
memory-efficient training |
0.8 | 1 | 2024 | Full Parameter Fine-tuning for Large Language Models with Limited Resources · ACL (1) 2024 |
Machine learning › Efficient and distributed learning
memory optimization |
0.8 | 1 | 2024 | Full Parameter Fine-tuning for Large Language Models with Limited Resources · ACL (1) 2024 |
Natural language and speech › Language models and text generation › in-context learning
demonstration retrieval |
0.7 | 1 | 2023 | Unified Demonstration Retriever for In-Context Learning · ACL (1) 2023 |
Natural language and speech › Language models and text generation
in-context learning |
0.7 | 1 | 2023 | Unified Demonstration Retriever for In-Context Learning · ACL (1) 2023 |
Natural language and speech › Language models and text generation › text generation
neural text generation |
0.6 | 1 | 2022 | CoNT: Contrastive Neural Text Generation · NeurIPS 2022 |
Information retrieval
retrieval models |
0.2 | 1 | 2023 | Unified Demonstration Retriever for In-Context Learning · ACL (1) 2023 |
Natural language and speech › Language models and text generation
text summarization |
0.2 | 1 | 2022 | CoNT: Contrastive Neural Text Generation · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
preference mining · 1.7triton · 0.9top-k attention · 0.9soft blocking · 0.9routing model · 0.9rejection sampling · 0.9multiple sampling · 0.9monte carlo tree search · 0.9hard blocking · 0.9gradient fusion · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CritiQ: Mining Data Quality Criteria from Human PreferencesabstractHonglin Guo, Kai Lv, Qipeng Guo, Tianyi Liang, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun, Kai Chen, Xipeng Qiu, Tao Gui. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Honglin Guo, Kai Lv 0001, Qipeng Guo, Tianyi Liang 0002, Zhiheng Xi, Demin Song, Qiuyinzhe Zhang, Yu Sun 0031, Kai Chen 0026, Xipeng Qiu, Tao Gui |
ACL (1) | 2 |
| 2025 | FastMCTS: A Simple Sampling Strategy for Data SynthesisabstractSynthetic high-quality multi-step reasoning data can significantly enhance the performance of large language models on various tasks. However, most existing methods rely on rejection sampling, which generates trajectories independently and suffers from inefficiency and imbalanced sampling across problems of varying difficulty. In this work, we introduce FastMCTS, an innovative data synthesis strategy inspired by Monte Carlo Tree Search. FastMCTS provides a more efficient sampling method for multi-step reasoning data, offering step-level evaluation signals and promoting balanced sampling across problems of different difficulty levels. Experiments on both English and Chinese reasoning datasets demonstrate that FastMCTS generates over 30% more correct reasoning paths compared to rejection sampling as the number of generated tokens scales up. Furthermore, under comparable synthetic data budgets, models trained on FastMCTS-generated data outperform those trained on rejection sampling data by 3.9% across multiple benchmarks. As a lightweight sampling strategy, FastMCTS offers a practical and efficient alternative for synthesizing high-quality reasoning data. Peiji Li, Kai Lv 0001, Yunfan Shao, Yichuan Ma, Linyang Li, Xiaoqing Zheng, Xipeng Qiu, Qipeng Guo |
ACL (1) | 2 |
| 2025 | Firewall Routing: Blocking Leads to Better Hybrid Inference for LLMsabstractThe rapid advancement of Large Language Models (LLMs) has significantly enhanced performance across various natural language processing (NLP) tasks, yet the high computational costs and latency associated with deploying such models continue to pose critical bottlenecks, limiting their broader applicability.To mitigate these challenges, we propose a dynamic hybrid inference framework, Firewall Routing, which efficiently selects between a strong and a weak LLMs based on the complexity of the query.A lightweight routing model is trained to optimize resource allocation by learning from response quality and preventing longtail queries, which are often too hard to solve by LLMs, from being routed to the stronger model.Moreover, our method incorporates multiple sampling to enhance query evaluation reliability while leveraging Hard Blocking and Soft Blocking to handle long-tail queries along with refining labels for model selection.Extensive experiments show our method outperforms existing routing strategies by up to 5.29% in APGR, demonstrating state-of-the-art performance across multiple benchmarks. Runyu Peng, Yunhua Zhou, Kai Lv 0001, Yang Gao 0042, Qipeng Guo, Xipeng Qiu |
EMNLP | 3 |
| 2025 | ReAttention: Training-Free Infinite Context with Finite Attention ScopeabstractThe long-context capability of the Large Language Models (LLM) has made significant breakthroughs, but \textit{the maximum supported context length in length extrapolation} remains a critical bottleneck limiting their practical applications. The constraint of context length in LLMs arises from the self-attention mechanism, which cannot effectively and efficiently capture the semantic relationships within infinitely long contexts via the limited pre-trained positional information and attention scope. In this work, we propose \textbf{ReAttention}, a training-free approach enabling LLM based on the self-attention mechanism to support an infinite context with a finite attention scope under sufficient memory resources. ReAttention performs the position-agnostic top-$k$ attention before the ordinary position-aware self-attention, freeing LLMs from the length extrapolation issue. We validate the performance of ReAttention on the LongBench, L-Eval, and InfiniteBench and demonstrate that it is on par with traditional methods. Furthermore, we also apply ReAttention on mainstream LLMs, including LLaMA3.1-8B and Mistral-v0.3-7B, enabling them to support context lengths of at least 1M and even expanding the context length of LLaMA3.2-3B-chat by 128$\times$ to 4M without any further training in Needle-In-A-Haystack tests. We also improve the efficiency of ReAttention with Triton and achieve an efficient extrapolation without additional overhead. The code is available at \url{https://github.com/OpenMOSS/ReAttention}. Ruixiao Li, Zhigeng Liu, Qipeng Guo, Yuerong Song, Kai Lv 0001, Hang Yan 0001, Linlin Li 0001, Qun Liu 0001, Xipeng Qiu |
ICLR | 6 |
| 2024 | Full Parameter Fine-tuning for Large Language Models with Limited ResourcesabstractLarge Language Models (LLMs) have revolutionized Natural Language Processing (NLP) but demand massive GPU resources for training.Lowering the threshold for LLMs training would encourage greater participation from researchers, benefiting both academia and society.While existing approaches have focused on parameter-efficient fine-tuning, which tunes or adds a small number of parameters, few have addressed the challenge of tuning the full parameters of LLMs with limited resources.In this work, we propose a new optimizer, LOw-Memory Optimization (LOMO), which fuses the gradient computation and the parameter update in one step to reduce memory usage.By integrating LOMO with existing memory saving techniques, we reduce memory usage to 10.8% compared to the standard approach (DeepSpeed solution).Consequently, our approach enables the full parameter fine-tuning of a 65B model on a single machine with 8×RTX 3090, each with 24GB memory. 1 Kai Lv 0001, Yuqing Yang 0004, Tengxiao Liu, Qipeng Guo, Xipeng Qiu |
ACL (1) | 1 |
| 2023 | Unified Demonstration Retriever for In-Context LearningabstractXiaonan Li, Kai Lv, Hang Yan, Tianyang Lin, Wei Zhu, Yuan Ni, Guotong Xie, Xiaoling Wang, Xipeng Qiu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Kai Lv 0001, Hang Yan 0001, Tianyang Lin, Wei Zhu 0016, Yuan Ni, Guo Tong Xie, Xiaoling Wang 0004, Xipeng Qiu |
ACL (1) | 2 |
| 2022 | CoNT: Contrastive Neural Text GenerationabstractRecently, contrastive learning attracts increasing interests in neural text generation as a new solution to alleviate the exposure bias problem. It introduces a sequence-level training signal which is crucial to generation tasks that always rely on auto-regressive decoding. However, previous methods using contrastive learning in neural text generation usually lead to inferior performance. In this paper, we analyse the underlying reasons and propose a new Contrastive Neural Text generation framework, CoNT. CoNT addresses bottlenecks that prevent contrastive learning from being widely adopted in generation tasks from three aspects -- the construction of contrastive examples, the choice of the contrastive loss, and the strategy in decoding. We validate CoNT on five generation tasks with ten benchmarks, including machine translation, summarization, code comment generation, data-to-text generation and commonsense generation. Experimental results show that CoNT clearly outperforms its baseline on all the ten benchmarks with a convincing margin. Especially, CoNT surpasses previous the most competitive contrastive learning method for text generation, by 1.50 BLEU on machine translation and 1.77 ROUGE-1 on summarization, respectively. It achieves new state-of-the-art on summarization, code comment generation (without external data) and data-to-text generation. Chenxin An, Jiangtao Feng, Kai Lv 0001, Lingpeng Kong, Xipeng Qiu, Xuanjing Huang 0001 |
NeurIPS | 3 |