EDBT 2026 Demo / reviewers in the wild / expert
Qingrong Xia
dblp:185/0855
· DBLP profile ↗
19ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0003-4578-119XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Branch Self-Drafting for LLM Inference AccelerationabstractThe autoregressive decoding paradigm endows large language models (LLMs) with superior language generation capabilities; however, its step-by-step decoding process inherently limits decoding speed. To mitigate these constraints, the prevalent “draft and validation” strategy enables parallel validation of candidate drafts, allowing LLMs to decode multiple tokens simultaneously during one model forward propagation. However, existing methodologies for obtaining drafts often incur additional overhead in communication or training process, or statistical biases from the corpus. To this end, we propose an innovative draft generation and maintenance approach that leverages the capabilities of LLM itself. Specifically, we extend the autoregressive decoding paradigm to a multi-branch drafting procedure, which can efficiently generate draft sequences without any additional models or training process, while preserving the quality of the generated content by maintaining LLM parameters. Experiments across various open-source benchmarks show that our method generates 2.0 to 3.2 tokens per forward step and achieves around 2 times improvement of end-to-end throughput compared to the autoregressive decoding strategy. Zipeng Gao, Qingrong Xia, Tong Xu 0001, Xinyu Duan, Zhi Zheng 0008, Zhefeng Wang 0001, Enhong Chen |
AAAI | 2 |
| 2025 | Accurate KV Cache Quantization with Outlier Tokens TracingabstractThe impressive capabilities of Large Language Models (LLMs) come at the cost of substantial computational resources during deployment. While KV Cache can significantly reduce recomputation during inference, it also introduces additional memory overhead. KV Cache quantization presents a promising solution, striking a good balance between memory usage and accuracy. Previous research has shown that the Keys are distributed by channel, while the Values are distributed by token. Consequently, the common practice is to apply channel-wise quantization to the Keys and token-wise quantization to the Values. However, our further investigation reveals that a small subset of unusual tokens exhibit unique characteristics that deviate from this pattern, which can substantially impact quantization accuracy. To address this, we develop a simple yet effective method to identify these tokens accurately during the decoding process and exclude them from quantization as outlier tokens, significantly improving overall accuracy. Extensive experiments show that our method achieves significant accuracy improvements under 2-bit quantization and can deliver a 6.4 times reduction in memory usage and a 2.3 times increase in throughput. Yi Su 0006, Yuechi Zhou, Quantong Qiu, Juntao Li 0005, Qingrong Xia, Ping Li 0016, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
ACL (1) | 5 |
| 2025 | Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional VerificationabstractRecent works have revealed the great potential of speculative decoding in accelerating the autoregressive generation process of large language models.The success of these methods relies on the alignment between draft candidates and the sampled outputs of the target model.Existing methods mainly achieve draft-target alignment with training-based methods, e.g., EAGLE, Medusa, involving considerable training costs.In this paper, we present a trainingfree alignment-augmented speculative decoding algorithm.We propose alignment sampling, which leverages output distribution obtained in the prefilling phase to provide more aligned draft candidates.To further benefit from highquality but non-aligned draft candidates, we also introduce a simple yet effective flexible verification strategy.Through an adaptive probability threshold, our approach can improve generation accuracy while further improving inference efficiency.Experiments on 8 datasets (including question answering, summarization and code completion tasks) show that our approach increases the average generation score by 3.3 points for the LLaMA3 model.Our method achieves a mean acceptance length up to 2.39 and speed up generation by 2.23×. Zhenxu Tian, Juntao Li 0005, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005 |
EMNLP | 4 |
| 2025 | Beware of Calibration Data for Pruning Large Language ModelsabstractAs large language models (LLMs) are widely applied across various fields, model
compression has become increasingly crucial for reducing costs and improving
inference efficiency. Post-training pruning is a promising method that does not
require resource-intensive iterative training and only needs a small amount of
calibration data to assess the importance of parameters. Recent research has enhanced post-training pruning from different aspects but few of them systematically
explore the effects of calibration data, and it is unclear if there exist better calibration data construction strategies. We fill this blank and surprisingly observe that
calibration data is also crucial to post-training pruning, especially for high sparsity. Through controlled experiments on important influence factors of calibration
data, including the pruning settings, the amount of data, and its similarity with
pre-training data, we observe that a small size of data is adequate, and more similar data to its pre-training stage can yield better performance. As pre-training data
is usually inaccessible for advanced LLMs, we further provide a self-generating
calibration data synthesis strategy to construct feasible calibration data. Experimental results on recent strong open-source LLMs (e.g., DCLM, and LLaMA-3)
show that the proposed strategy can enhance the performance of strong pruning
methods (e.g., Wanda, DSnoT, OWL) by a large margin (up to 2.68%). Yixin Ji, Yang Xiang 0003, Juntao Li 0005, Qingrong Xia, Ping Li 0016, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
ICLR | 4 |
| 2025 | Taming the Titans: A Survey of Efficient LLM Inference ServingabstractLarge Language Models (LLMs) for Generative AI have achieved remarkable progress, evolving into sophisticated and versatile tools widely adopted across various domains and applications. However, the substantial memory overhead caused by their vast number of parameters, combined with the high computational demands of the attention mechanism, poses significant challenges in achieving low latency and high throughput for LLM inference services. Recent advancements, driven by groundbreaking research, have significantly accelerated progress in this field. This paper provides a comprehensive survey of these methods, covering fundamental instance-level approaches, in-depth cluster-level strategies, and emerging scenarios. At the instance level, we review model placement, request scheduling, decoding length prediction, storage management, and the disaggregation paradigm. At the cluster level, we explore GPU cluster deployment, multi-instance load balancing, and cloud service solutions. Additionally, we discuss specific tasks, modules, and auxiliary methods in emerging scenarios. Finally, we outline potential research directions to further advance the field of LLM inference serving. Ranran Zhen, Juntao Li 0005, Yixin Ji, Zhenlin Yang, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005 |
INLG | 6 |
| 2025 | OPT-Tree: Speculative Decoding with Adaptive Draft Tree StructureabstractAbstract Autoregressive language models demonstrate excellent performance in various scenarios. However, the inference efficiency is limited by its one-step-one-word generation mode, which has become a pressing problem recently as the models become increasingly larger. Speculative decoding employs a “draft and then verify” mechanism to allow multiple tokens to be generated in one step, realizing lossless acceleration. Existing methods mainly adopt fixed heuristic draft structures, which do not adapt to different situations to maximize the acceptance length during verification. To alleviate this dilemma, we propose OPT-Tree, an algorithm to construct adaptive and scalable draft trees, which can be applied to any autoregressive draft model. It searches the optimal tree structure that maximizes the mathematical expectation of the acceptance length in each decoding step. Experimental results reveal that OPT-Tree outperforms the existing draft structures and achieves a speed-up ratio of up to 3.2 compared with autoregressive decoding. If the draft model is powerful enough and the node budget is sufficient, it can generate more than ten tokens in a single step. Our code is available at https://github.com/Jikai0Wang/OPT-Tree. Yi Su 0006, Juntao Li 0005, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | When and How to Grow? On Efficient Pre-training via Model Growth
Juntao Li 0005, Min Zhang 0005, Zechang Li, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai |
ACML | 5 |
| 2024 | Are Bert Family Good Instruction Followers? A Study on Their Potential And LimitationsabstractLanguage modeling at scale has proven very effective and brought unprecedented success to natural language models. Many typical representatives, especially decoder-only models, e.g., BLOOM and LLaMA, and encoder-decoder models, e.g., Flan-T5 and AlexaTM, have exhibited incredible instruction-following capabilities while keeping strong task completion ability. These large language models can achieve superior performance in various tasks and even yield emergent capabilities, e.g., reasoning and universal generalization. Though the above two paradigms are mainstream and well explored, the potential of the BERT family, which are encoder-only based models and have ever been one of the most representative pre-trained models, also deserves attention, at least should be discussed. In this work, we adopt XML-R to explore the effectiveness of the BERT family for instruction following and zero-shot learning. We first design a simple yet effective strategy to utilize the encoder-only models for generation tasks and then conduct multi-task instruction tuning. Experimental results demonstrate that our fine-tuned model, Instruct-XMLR, outperforms Bloomz on all evaluation tasks and achieves comparable performance with mT0 on most tasks. Surprisingly, Instruct-XMLR also possesses strong task and language generalization abilities, indicating that Instruct-XMLR can also serve as a good instruction follower and zero-shot learner. Besides, Instruct-XMLR can accelerate decoding due to its non-autoregressive generation manner, achieving around 3 times speedup compared with current autoregressive large language models. Although we also witnessed several limitations through our experiments, such as the performance decline in long-generation tasks and the shortcoming of length prediction, Instruct-XMLR can still become a good member of the family of current large language models. Yisheng Xiao, Juntao Li 0005, Zechen Sun, Zechang Li, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
ICLR | 5 |
| 2024 | A Survey on Arabic Named Entity Recognition: Past, Recent Advances, and Future TrendsabstractAs more and more Arabic texts emerged on the Internet, extracting important information from these Arabic texts is especially useful. As a fundamental technology, Named entity recognition (NER) serves as the core component in information extraction technology, while also playing a critical role in many other Natural Language Processing (NLP) systems, such as question answering and knowledge graph building. In this paper, we provide a comprehensive review of the development of Arabic NER, especially the recent advances in deep learning and pre-trained language model. Specifically, we first introduce the background of Arabic NER, including the characteristics of Arabic and existing resources for Arabic NER. Then, we systematically review the development of Arabic NER methods. Traditional Arabic NER systems focus on feature engineering and designing domain-specific rules. In recent years, deep learning methods achieve significant progress by representing texts via continuous vector representations. With the growth of pre-trained language model, Arabic NER yields better performance. Finally, we conclude the method gap between Arabic NER and NER methods from other languages, which helps outline future directions for Arabic NER. Xiaoye Qu, Yingjie Gu, Qingrong Xia, Zechang Li, Zhefeng Wang 0001, Baoxing Huai |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Semantic Role Labeling as Dependency Parsing: Exploring Latent Tree Structures inside ArgumentsabstractSemantic role labeling (SRL) is a fundamental yet challenging task in the NLP community. Recent works of SRL mainly fall into two lines: 1) BIO-based; 2) span-based. Despite ubiquity, they share some intrinsic drawbacks of not considering internal argument structures, potentially hindering the model’s expressiveness. The key challenge is arguments are flat structures, and there are no determined subtree realizations for words inside arguments. To remedy this, in this paper, we propose to regard flat argument spans as latent subtrees, accordingly reducing SRL to a tree parsing task. In particular, we equip our formulation with a novel span-constrained TreeCRF to make tree structures span-aware and further extend it to the second-order case. We conduct extensive experiments on CoNLL05 and CoNLL12 benchmarks. Results reveal that our methods perform favorably better than all previous syntax-agnostic works, achieving new state-of-the-art under both end-to-end and w/ gold predicates settings. Yu Zhang 0092, Qingrong Xia, Shilin Zhou 0002, Guohong Fu, Min Zhang 0005 |
COLING | 2 |
| 2022 | Fast and Accurate End-to-End Span-based Semantic Role Labeling as Word-based Graph ParsingabstractThis paper proposes to cast end-to-end span-based SRL as a word-based graph parsing task. The major challenge is how to represent spans at the word level. Borrowing ideas from research on Chinese word segmentation and named entity recognition, we propose and compare four different schemata of graph representation, i.e., BES, BE, BIES, and BII, among which we find that the BES schema performs the best. We further gain interesting insights through detailed analysis. Moreover, we propose a simple constrained Viterbi procedure to ensure the legality of the output graph according to the constraints of the SRL structure. We conduct experiments on two widely used benchmark datasets, i.e., CoNLL05 and CoNLL12. Results show that our word-based graph parsing approach achieves consistently better performance than previous results, under all settings of end-to-end and predicate-given, without and with pre-trained language models (PLMs). More importantly, our model can parse 669/252 sentences per second, without and with PLMs respectively. Shilin Zhou 0002, Qingrong Xia, Zhenghua Li, Yu Zhang 0092, Yu Hong 0001, Min Zhang 0005 |
COLING | 2 |
| 2022 | MuCPAD: A Multi-Domain Chinese Predicate-Argument DatasetabstractYahui Liu, Haoping Yang, Chen Gong, Qingrong Xia, Zhenghua Li, Min Zhang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Haoping Yang, Chen Gong 0004, Qingrong Xia, Zhenghua Li, Min Zhang 0005 |
NAACL-HLT | 4 |
| 2021 | A Unified Span-Based Approach for Opinion Mining with Syntactic ConstituentsabstractQingrong Xia, Bo Zhang, Rui Wang, Zhenghua Li, Yue Zhang, Fei Huang, Luo Si, Min Zhang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Qingrong Xia, Bo Zhang 0071, Rui Wang 0005, Zhenghua Li, Yue Zhang 0004, Fei Huang 0002, Luo Si, Min Zhang 0005 |
NAACL-HLT | 1 |
| 2021 | Emotion Classification with Explicit and Implicit Syntactic Information
Qingrong Xia, Xiabing Zhou, Wenliang Chen, Min Zhang 0005 |
NLPCC (1) | 2 |
| 2020 | Semantic Role Labeling with Heterogeneous Syntactic KnowledgeabstractRecently, due to the interplay between syntax and semantics, incorporating syntactic knowledge into neural semantic role labeling (SRL) has achieved much attention.Most of the previous syntax-aware SRL works focus on explicitly modeling homogeneous syntactic knowledge over tree outputs.In this work, we propose to encode heterogeneous syntactic knowledge for SRL from both explicit and implicit representations.First, we introduce graph convolutional networks to explicitly encode multiple heterogeneous dependency parse trees.Second, we extract the implicit syntactic representations from syntactic parser trained with heterogeneous treebanks.Finally, we inject the two types of heterogeneous syntax-aware representations into the base SRL model as extra inputs.We conduct experiments on two widely-used benchmark datasets, i.e., Chinese Proposition Bank 1.0 and English CoNLL-2005 dataset.Experimental results show that incorporating heterogeneous syntactic knowledge brings significant improvements over strong baselines.We further conduct detailed analysis to gain insights on the usefulness of heterogeneous (vs.homogeneous) syntactic knowledge and the effectiveness of our proposed approaches for modeling such knowledge. Qingrong Xia, Rui Wang 0005, Zhenghua Li, Yue Zhang 0004, Min Zhang 0005 |
COLING | 1 |
| 2020 | Hierarchical LSTM with char-subword-word tree-structure representation for Chinese named entity recognition
Chen Gong 0004, Zhenghua Li, Qingrong Xia, Wenliang Chen, Min Zhang 0005 |
Sci. China Inf. Sci. | 3 |
| 2019 | Syntax-Aware Neural Semantic Role LabelingabstractSemantic role labeling (SRL), also known as shallow semantic parsing, is an important yet challenging task in NLP. Motivated by the close correlation between syntactic and semantic structures, traditional discrete-feature-based SRL approaches make heavy use of syntactic features. In contrast, deep-neural-network-based approaches usually encode the input sentence as a word sequence without considering the syntactic structures. In this work, we investigate several previous approaches for encoding syntactic trees, and make a thorough study on whether extra syntax-aware representations are beneficial for neural SRL models. Experiments on the benchmark CoNLL-2005 dataset show that syntax-aware SRL approaches can effectively improve performance over a strong baseline with external word representations from ELMo. With the extra syntax-aware representations, our approaches achieve new state-of-the-art 85.6 F1 (single model) and 86.6 F1 (ensemble) on the test data, outperforming the corresponding strong baselines with ELMo by 0.8 and 1.0, respectively. Detailed error analysis are conducted to gain more insights on the investigated approaches. Qingrong Xia, Zhenghua Li, Min Zhang 0005, Meishan Zhang, Guohong Fu, Rui Wang 0005, Luo Si |
AAAI | 1 |
| 2019 | A Syntax-aware Multi-task Learning Framework for Chinese Semantic Role LabelingabstractQingrong Xia, Zhenghua Li, Min Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Qingrong Xia, Zhenghua Li, Min Zhang 0005 |
EMNLP/IJCNLP (1) | 1 |
| 2017 | Dependency Parsing with Partial Annotations: An Empirical ComparisonabstractThis paper describes and compares two straightforward approaches for dependency parsing with partial annotations (PA). The first approach is based on a forest-based training objective for two CRF parsers, i.e., a biaffine neural network graph-based parser (Biaffine) and a traditional log-linear graph-based parser (LLGPar). The second approach is based on the idea of constrained decoding for three parsers, i.e., a traditional linear graph-based parser (LGPar), a globally normalized neural network transition-based parser (GN3Par) and a traditional linear transition-based parser (LTPar). For the test phase, constrained decoding is also used for completing partial trees. We conduct experiments on Penn Treebank under three different settings for simulating PA, i.e., random, most uncertain, and divergent outputs from the five parsers. The results show that LLGPar is most effective in directly learning from PA, and other parsers can achieve best performance when PAs are completed into full trees by LLGPar. Yue Zhang 0004, Zhenghua Li, Jun Lang 0001, Qingrong Xia, Min Zhang 0005 |
IJCNLP(1) | 4 |