VLDB 2026 Research / reviewers in the wild / expert
Guoming Liu
dblp:288/8406
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Efficient and distributed learning · 63% Language models and text generation · 32% Representation and self-supervised learning · 5% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
inference acceleration |
1.9 | 2 | 2026 | Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios · AAAI 2026 Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding · EMNLP 2025 |
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding |
1.9 | 2 | 2026 | Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios · AAAI 2026 Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding · EMNLP 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
1.7 | 2 | 2025 | SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers · ACL (1) 2025 KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding · ACL (1) 2025 |
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression |
1.7 | 2 | 2025 | SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers · ACL (1) 2025 KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding · ACL (1) 2025 |
Information retrieval
cross-modal retrieval |
1.0 | 1 | 2026 | End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering · AAAI 2026 |
Natural language and speech › Language models and text generation › neural language model
bidirectional language model |
0.9 | 1 | 2025 | What Limits Bidirectional Model's Generative Capabilities? A Uni-Bi-Directional Mixture-of-Expert Method For Bidirectional Fine-tuning · ICML 2025 |
Machine learning › Efficient and distributed learning
KV cache management |
0.9 | 1 | 2025 | SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers · ACL (1) 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization |
0.9 | 1 | 2025 | XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model inference
long-context inference |
0.9 | 1 | 2025 | XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025 |
Natural language and speech › Language models and text generation › language modeling
long-context language modeling |
0.9 | 1 | 2025 | KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding · ACL (1) 2025 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling |
0.9 | 1 | 2025 | SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers · ACL (1) 2025 |
Machine learning › Efficient and distributed learning
model compression |
0.9 | 1 | 2025 | XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025 |
Natural language and speech › Language models and text generation › prompting
prompt compression |
0.9 | 1 | 2025 | DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression · ACL (1) 2025 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.9 | 1 | 2025 | XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025 |
Machine learning › Representation and self-supervised learning
text embedding |
0.9 | 1 | 2025 | What Limits Bidirectional Model's Generative Capabilities? A Uni-Bi-Directional Mixture-of-Expert Method For Bidirectional Fine-tuning · ICML 2025 |
Natural language and speech › Language models and text generation
large language model inference |
0.3 | 1 | 2026 | Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios · AAAI 2026 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
spoken question answering |
0.3 | 1 | 2026 | End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering · AAAI 2026 |
Natural language and speech › Language models and text generation
in-context learning |
0.3 | 1 | 2025 | Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding · EMNLP 2025 |
Natural language and speech › Language models and text generation
long context |
0.3 | 1 | 2025 | DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression · ACL (1) 2025 |
Machine learning › Efficient and distributed learning › inference efficiency
memory-efficient inference |
0.3 | 1 | 2025 | XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
retrieval-augmented generation · 2.0contrastive language-speech pretraining · 2.0speculative decoding · 1.0non-autoregressive generation · 1.0bidirectional attention · 1.0rotary position embedding · 0.9layer-wise cache balancing · 0.9latent space down-sampling · 0.9entropy-based compression · 0.9attention-aware compression · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question AnsweringabstractSignificant progress has been made in spoken question answering (SQA) in recent years. However, many existing methods, including large audio language models, struggle with processing long audio. Follow the success of retrieval augmented generation, a speech-related retriever shows promising in help preprocessing long-form speech. But the performance of existing speech-related retrievers is lacking. To address this challenge, we propose CLSR, an end-to-end contrastive language-speech retriever that efficiently extracts question-relevant segments from long audio recordings for downstream SQA task. Unlike conventional speech-text contrastive models, CLSR incorporates an intermediate step that converts acoustic features into text-like representations prior to alignment, thereby more effectively bridging the gap between modalities. Experimental results across four cross-modal retrieval datasets demonstrate that CLSR surpasses both end-to-end speech related retrievers and pipeline approaches combining speech recognition with text retrieval, providing a robust foundation for advancing practical long-form SQA applications. Jiliang Hu 0001, Zuchao Li, Baoyuan Qi, Guoming Liu, Ping Wang 0028 |
AAAI | 4 |
| 2026 | Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch ScenariosabstractSpeculative decoding accelerates LLM inference by utilizing otherwise idle computational resources during memory-to-chip data transfer. Current speculative decoding methods typically assume a considerable amount of available computing power, then generate a complex and massive draft tree using a small autoregressive language model to improve overall prediction accuracy. However, methods like batching have been widely applied in mainstream model inference systems as a superior alternative to speculative decoding, as they compress the available idle computing power. Therefore, performing speculative decoding with low verification resources and low scheduling costs has become an important research problem. We believe that more capable models that allow for parallel generation on draft sequences are what we truly need. Recognizing the fundamental nature of draft models to only generate sequences of limited length, we propose SpecFormer, a novel architecture that integrates unidirectional and bidirectional attention mechanisms. SpecFormer combines the autoregressive model’s ability to extract information from the entire input sequence with the parallel generation benefits of non-autoregressive models. This design eliminates the reliance on large prefix trees and achieves consistent acceleration, even in large-batch scenarios. Through lossless speculative decoding experiments across models of various scales, we demonstrate that SpecFormer sets a new standard for scaling LLM inference with lower training demands and reduced computational costs. Luohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi, Guoming Liu, Hai Zhao 0001 |
AAAI | 5 |
| 2025 | KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional EmbeddingabstractLarge language models (LLMs) based on Transformer Decoders have become the preferred choice for conversational generative AI.Despite the overall superiority of the Decoder architecture, the gradually increasing Key-Value (KV) cache during inference has emerged as a primary efficiency bottleneck, both in aspects of memory consumption and data transfer bandwidth limitations.To address these challenges, we propose a paradigm called KV-Latent.By down-sampling the Key-Value vector dimensions into a latent space, we can significantly reduce the KV Cache footprint and improve inference speed, only with a small amount of extra training, less than 1% of pretraining takes.Besides, we enhanced the stability of Rotary Positional Embedding applied on lower-dimensional vectors by modifying its frequency sampling mechanism, avoiding noise introduced by higher frequencies while retaining position attenuation.Our experiments, including both models with Grouped Query Attention and those without, have yielded satisfactory results.Finally, we conducted comparative experiments to study the impact of separately reducing Key and Value components on model's performance.Our approach allows for the construction of more efficient language model systems, and opens the new possibility on KV Cache saving and efficient LLMs.Our code is available at https://github.com/ShiLuohe/KV- Latent. Luohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi, Guoming Liu, Hai Zhao 0001 |
ACL (1) | 5 |
| 2025 | SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep LayersabstractZicong Tang, Shi Luohe, Zuchao Li, Baoyuan Qi, Liu Guoming, Lefei Zhang, Ping Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zicong Tang, Luohe Shi, Zuchao Li, Baoyuan Qi, Guoming Liu, Lefei Zhang, Ping Wang 0028 |
ACL (1) | 5 |
| 2025 | DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt CompressionabstractTask-agnostic prompt compression leverages the redundancy in natural language to reduce computational overhead and enhance information density within prompts, especially in longcontext scenarios.Existing methods predominantly rely on information entropy as the metric to compress lexical units, aiming to achieve minimal information loss.However, these approaches overlook two critical aspects: (i) the importance of attention-critical tokens at the algorithmic level, and (ii) shifts in information entropy during the compression process.Motivated by these challenges, we propose a dynamic attention-aware approach for taskagnostic prompt compression (DAC).This approach effectively integrates entropy and attention information, dynamically sensing entropy shifts during compression to achieve fine-grained prompt compression.Extensive experiments across various domains, including LongBench, GSM8K, and BBH, show that DAC consistently yields robust and substantial improvements across a diverse range of tasks and LLMs, offering compelling evidence of its efficacy. Zuchao Li, Hai Zhao 0001, Baoyuan Qi, Guoming Liu |
ACL (1) | 5 |
| 2025 | Faster In-Context Learning for LLMs via N-Gram Trie Speculative DecodingabstractAs a crucial method in prompt engineering, In-Context Learning (ICL) enhances the generalization and knowledge utilization capabilities of Large Language Models (LLMs) (Dong et al., 2024).However, the lengthy retrieved contexts and limited token throughput in autoregressive models significantly constrain reasoning speed.To address this challenge, we propose N-Gram Trie Speculative Decoding, a novel approach that leverages the overlap between context and model output.This method constructs an n-gram trie from the context to generate drafts, accelerating token generation for LLMs.We evaluate our approach on summarization, Retrieval-Augmented Generation (RAG), and contextbased Question Answering (QA) tasks.Experimental results on Vicuna-7B, Llama2-7B-Chat, and Llama3-8B-Instruct demonstrate substantial speed improvements without compromising accuracy.Compared with various strong baselines, our method achieves the highest mean speedup, showcasing its effectiveness and efficiency.Our implement code is available here: https://github.com/mrlife219/Ngram-Trie. Jinglin Chen, Qiwei Li 0002, Zuchao Li, Baoyuan Qi, Guoming Liu, Haojun Ai, Hai Zhao 0001, Ping Wang 0028 |
EMNLP | 5 |
| 2025 | XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer CompressionabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks.However, their extensive memory requirements, particularly due to KV cache growth during long-text understanding and generation, present significant challenges for deployment in resourceconstrained environments.Quantization has emerged as a promising solution to reduce memory consumption while preserving historical information.We propose XQuant, a training-free and plug-and-play framework that achieves ultra-low equivalent bit-width KV cache quantization.XQuant introduces two key innovations: a computationally negligible data-free calibration method and cross-layer KV cache compression, enabling quantization to sub-1.4 bits.Extensive experiments on TruthfulQA and LongBench demonstrate that XQuant outperforms state-of-the-art methods (e.g., KIVI-2bit and AsymKV-1.5bit)by achieving lower bit-width while maintaining superior performance, establishing a better trade-off between memory efficiency and model accuracy.The source code is available at https: //github.com/brinenick511/XQuant.KeyCache[l][0], 14 KeyCache[l][1], 15 KeyCache[l][2] 16 else 17 DequantizedKey ← Dequantize 18 KeyCache[l -1][0], 19 KeyCache[l -1][1], 20 KeyCache[l][2] 21 if l < vm or l mod 2 == 0 then 22 DequantizedValue ← Dequantize 23 ValueCache[l][0], 24 ValueCache[l][1], 25 ValueCache[l][2] 26 else 27 DequantizedValue ← Dequantize 28 ValueCache[l -1][0], 29 ValueCache[l -1][1], 30 ValueCache[l][2] Haoqi Yang 0001, Yao Yao 0008, Zuchao Li, Baoyuan Qi, Guoming Liu, Hai Zhao 0001 |
EMNLP | 5 |
| 2025 | What Limits Bidirectional Model's Generative Capabilities? A Uni-Bi-Directional Mixture-of-Expert Method For Bidirectional Fine-tuningabstractLarge Language Models (LLMs) excel in generation tasks, yet their causal attention mechanisms limit performance in embedding tasks. While bidirectional modeling may enhance embeddings, naively fine-tuning unidirectional models bidirectionally severely degrades generative performance. To investigate this trade-off, we analyze attention weights as dependence indicators and find that bidirectional fine-tuning increases subsequent dependence, impairing unidirectional generation. Through systematic Transformer module evaluations, we discover the FFN layer is least affected by such dependence. Leveraging this discovery, we propose UBMoE-LLM, a novel Uni-Bi-directional Mixture-of-Experts LLM, which integrates the original unidirectional FFN with a bidirectionally fine-tuned FFN via unsupervised contrastive learning. This MoE-based approach enhances embedding performance while preserving robust generation. Extensive experiments across diverse datasets and model scales validate our attention dependence metric and demonstrate UBMoE-LLM’s superior generative quality and reduced hallucination. Code is available at: https://github.com/heiyonghua/ubmoe_llm. Zuchao Li, Yonghua Hei, Qiwei Li 0002, Lefei Zhang, Ping Wang 0028, Hai Zhao 0001, Baoyuan Qi, Guoming Liu |
ICML | 8 |
| 2024 | Cross-domain NER in the data-poor scenarios for human mobility knowledge
Fusheng Jin, Mengnan Chen, Guoming Liu, He Pang, Ye Yuan 0001 |
GeoInformatica | 4 |