VLDB 2026 Research / reviewers in the wild / expert
Zhitao Li 0002
dblp:71/7318-2
· DBLP profile ↗
12ranked-venue papers
0as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model CompressionabstractRecent studies have demonstrated that many layers are functionally redundant in large language models (LLMs), enabling model compression by removing these layers to reduce inference cost.While such approaches can improve efficiency, indiscriminate layer pruning often results in significant performance degradation.In this paper, we propose GRASP (Gradient-based Retention of Adaptive Singular Parameters), a novel compression framework that mitigates this issue by preserving sensitivity-aware singular values.Unlike direct layer pruning, GRASP leverages gradient-based attribution on a small calibration dataset to adaptively identify and retain critical singular components.By replacing redundant layers with only a minimal set of parameters, GRASP achieves efficient compression while maintaining strong performance with minimal overhead.Experiments across multiple LLMs show that GRASP consistently outperforms existing compression methods, achieving 90% of the original model's performance under 20% compression ratio.The source code is available at https://github.com/LyoAI/GRASP. Kainan Liu, Yong Zhang 0058, Ning Cheng 0001, Zhitao Li 0002, Jing Xiao 0006 |
EMNLP | 4 |
| 2025 | Self-Enhanced Reasoning Training: Activating Latent Reasoning in Small Models for Enhanced Reasoning DistillationabstractThe rapid advancement of large language models (LLMs) has significantly enhanced their reasoning abilities, enabling increasingly complex tasks. However, these capabilities often diminish in smaller, more computationally efficient models like GPT-2. Recent research shows that reasoning distillation can help small models acquire reasoning capabilities, but most existing methods focus primarily on improving teacher-generated reasoning paths. Our observations reveal that small models can generate high-quality reasoning paths during sampling, even without chain-of-thought prompting, though these paths are often latent due to their low probability under standard decoding strategies. To address this, we propose Self-Enhanced Reasoning Training (SERT), which activates and leverages latent reasoning capabilities in small models through self-training on filtered, self-generated reasoning paths under zero-shot conditions. Experiments using OpenAI’s GPT-3.5 as the teacher model and GPT-2 models as the student models demonstrate that SERT enhances the reasoning abilities of small models, improving their performance in reasoning distillation. Yong Zhang 0058, Zhitao Li 0002, Ming Li 0010, Ning Cheng 0001, Minchuan Chen, Tao Wei 0003, Jun Ma 0018, Jing Xiao 0006 |
ICASSP | 3 |
| 2024 | Superfiltering: Weak-to-Strong Data Filtering for Fast Instruction-TuningabstractMing Li, Yong Zhang, Shwai He, Zhitao Li, Hongyu Zhao, Jianzong Wang, Ning Cheng, Tianyi Zhou. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Ming Li 0010, Yong Zhang 0058, Shwai He, Zhitao Li 0002, Jianzong Wang, Ning Cheng 0001, Tianyi Zhou 0001 |
ACL (1) | 4 |
| 2024 | Leveraging Biases in Large Language Models: "bias-kNN" for Effective Few-Shot LearningabstractLarge Language Models (LLMs) have shown significant promise in various applications, including zero-shot and few-shot learning. However, their performance can be hampered by inherent biases. Instead of traditionally sought methods that aim to minimize or correct these biases, this study introduces a novel methodology named "bias-kNN". This approach capitalizes on the biased outputs, harnessing them as primary features for kNN and supplementing with gold labels. Our comprehensive evaluations, spanning diverse domain text classification datasets and different GPT-2 model sizes, indicate the adaptability and efficacy of the "bias-kNN" method. Remarkably, this approach not only outperforms conventional in-context learning in few-shot scenarios but also demonstrates robustness across a spectrum of samples, templates and verbalizers. This study, therefore, presents a unique perspective on harnessing biases, transforming them into assets for enhanced model performance. Yong Zhang 0058, Hanzhang Li, Zhitao Li 0002, Ning Cheng 0001, Ming Li 0010, Jing Xiao 0006, Jianzong Wang |
ICASSP | 3 |
| 2024 | QLSC: A Query Latent Semantic Calibrator for Robust Extractive Question AnsweringabstractExtractive Question Answering (EQA) in Machine Reading Comprehension (MRC) often faces the challenge of dealing with semantically identical but format-variant inputs. Our work introduces a novel approach, called the "Query Latent Semantic Calibrator (QLSC)", designed as an auxiliary module for existing MRC models. We propose a unique scaling strategy to capture latent semantic center features of queries. These features are then seamlessly integrated into traditional query and passage embeddings using an attention mechanism. By deepening the comprehension of the semantic queries-passage relationship, our approach diminishes sensitivity to variations in text format and boosts the model’s capability in pinpointing accurate answers. Experimental results on robust Question-Answer datasets confirm that our approach effectively handles format-variant but semantically identical queries, highlighting the effectiveness and adaptability of our proposed method. Sheng Ouyang, Jianzong Wang, Yong Zhang 0058, Zhitao Li 0002, Ziqi Liang, Xulong Zhang 0001, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 4 |
| 2024 | From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction TuningabstractMing Li, Yong Zhang, Zhitao Li, Jiuhai Chen, Lichang Chen, Ning Cheng, Jianzong Wang, Tianyi Zhou, Jing Xiao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Ming Li 0010, Yong Zhang 0058, Zhitao Li 0002, Jiuhai Chen, Lichang Chen, Ning Cheng 0001, Jianzong Wang, Tianyi Zhou 0001, Jing Xiao 0006 |
NAACL-HLT | 3 |
| 2023 | On the Calibration and Uncertainty with Pólya-Gamma Augmentation for Dialog Retrieval ModelsabstractDeep neural retrieval models have amply demonstrated their power but estimating the reliability of their predictions remains challenging. Most dialog response retrieval models output a single score for a response on how relevant it is to a given question. However, the bad calibration of deep neural network results in various uncertainty for the single score such that the unreliable predictions always misinform user decisions. To investigate these issues, we present an efficient calibration and uncertainty estimation framework PG-DRR for dialog response retrieval models which adds a Gaussian Process layer to a deterministic deep neural network and recovers conjugacy for tractable posterior inference by Pólya-Gamma augmentation. Finally, PG-DRR achieves the lowest empirical calibration error (ECE) in the in-domain datasets and the distributional shift task while keeping R10@1 and MAP performance. Shijing Si, Jianzong Wang, Ning Cheng 0001, Zhitao Li 0002, Jing Xiao 0006 |
AAAI | 5 |
| 2023 | PRCA: Fitting Black-Box Large Language Models for Retrieval Question Answering via Pluggable Reward-Driven Contextual AdapterabstractThe Retrieval Question Answering (ReQA) task employs the retrieval-augmented framework, composed of a retriever and generator.The generator formulates the answer based on the documents retrieved by the retriever.Incorporating Large Language Models (LLMs) as generators is beneficial due to their advanced QA capabilities, but they are typically too large to be fine-tuned with budget constraints while some of them are only accessible via APIs.To tackle this issue and further improve ReQA performance, we propose a trainable Pluggable Reward-Driven Contextual Adapter (PRCA), keeping the generator as a black box.Positioned between the retriever and generator in a Pluggable manner, PRCA refines the retrieved information by operating in a tokenautoregressive strategy via maximizing rewards of the reinforcement learning phase.Our experiments validate PRCA's effectiveness in enhancing ReQA performance on three datasets by up to 20% improvement to fit black-box LLMs into existing frameworks, demonstrating its considerable potential in the LLMs era. Zhitao Li 0002, Yong Zhang 0058, Jianzong Wang, Ning Cheng 0001, Ming Li 0010, Jing Xiao 0006 |
EMNLP | 2 |
| 2023 | Efficient Uncertainty Estimation with Gaussian Process for Reliable Dialog Response RetrievalabstractDeep neural networks have achieved remarkable performance in retrieval-based dialogue systems, but they are shown to be ill calibrated. Though basic calibration methods like Monte Carlo Dropout and Ensemble can calibrate well, these methods are time-consuming in the training or inference stages. To tackle these challenges, we propose an efficient uncertainty calibration framework GPF-BERT for BERT-based conversational search, which employs a Gaussian Process layer and the focal loss on top of the BERT architecture to achieve a high-quality neural ranker. Extensive experiments are conducted to verify the effectiveness of our method. In comparison with basic calibration methods, GPF-BERT achieves the lowest empirical calibration error (ECE) in three in-domain datasets and the distributional shift tasks, while yielding the highest R10@1 and MAP performance on most cases. In terms of time consumption, our GPF-BERT has an 8× speedup. Zhitao Li 0002, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 2 |
| 2023 | Boosting Chinese ASR Error Correction with Dynamic Error Scaling Mechanism
Yong Zhang 0058, Hanzhang Li, Jianzong Wang, Zhitao Li 0002, Sheng Ouyang, Ning Cheng 0001, Jing Xiao 0006 |
INTERSPEECH | 5 |
| 2023 | Prompt Guided Copy Mechanism for Conversational Question AnsweringabstractConversational Question Answering (CQA) is a challenging task that aims to generate natural answers for conversational flow questions.In this paper, we propose a pluggable approach for extractive methods that introduces a novel prompt-guided copy mechanism to improve the fluency and appropriateness of the extracted answers.Our approach uses prompts to link questions to answers and employs attention to guide the copy mechanism to verify the naturalness of extracted answers, making necessary edits to ensure that the answers are fluent and appropriate.The three prompts, including a question-rationale relationship prompt, a question description prompt, and a conversation history prompt, enhance the copy mechanism's performance.Our experiments demonstrate that this approach effectively promotes the generation of natural answers and achieves good results in the CoQA challenge. Yong Zhang 0058, Zhitao Li 0002, Jianzong Wang, Yiming Gao 0010, Ning Cheng 0001, Fengying Yu, Jing Xiao 0006 |
INTERSPEECH | 2 |
| 2022 | Self-Attention for Incomplete Utterance RewritingabstractIncomplete utterance rewriting (IUR) has recently become an essential task in NLP, aiming to complement the incomplete utterance with sufficient context information for comprehension. In this paper, we propose a novel method by directly extracting the coreference and omission relationship from the self-attention weight matrix of the transformer in-stead of word embeddings and edit the original text accordingly to generate the complete utterance. Benefiting from the rich information in the self-attention weight matrix, our method achieved competitive results on public IUR datasets. Yong Zhang 0058, Zhitao Li 0002, Jianzong Wang, Ning Cheng 0001, Jing Xiao 0006 |
ICASSP | 2 |