Guoming Liu

dblp:288/8406 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Efficient and distributed learning · 63% Language models and text generation · 32% Representation and self-supervised learning · 5%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference acceleration
1.922026
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios · AAAI 2026
Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding · EMNLP 2025
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
1.922026
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios · AAAI 2026
Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding · EMNLP 2025
Machine learning › Efficient and distributed learning
inference efficiency
1.722025
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers · ACL (1) 2025
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding · ACL (1) 2025
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression
1.722025
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers · ACL (1) 2025
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding · ACL (1) 2025
Information retrieval
cross-modal retrieval
1.012026
End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering · AAAI 2026
Natural language and speech › Language models and text generation › neural language model
bidirectional language model
0.912025
What Limits Bidirectional Model's Generative Capabilities? A Uni-Bi-Directional Mixture-of-Expert Method For Bidirectional Fine-tuning · ICML 2025
Machine learning › Efficient and distributed learning
KV cache management
0.912025
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers · ACL (1) 2025
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization
0.912025
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025
Natural language and speech › Language models and text generation › large language model inference
long-context inference
0.912025
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025
Natural language and speech › Language models and text generation › language modeling
long-context language modeling
0.912025
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding · ACL (1) 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling
0.912025
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers · ACL (1) 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025
Natural language and speech › Language models and text generation › prompting
prompt compression
0.912025
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression · ACL (1) 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.912025
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025
Machine learning › Representation and self-supervised learning
text embedding
0.912025
What Limits Bidirectional Model's Generative Capabilities? A Uni-Bi-Directional Mixture-of-Expert Method For Bidirectional Fine-tuning · ICML 2025
Natural language and speech › Language models and text generation
large language model inference
0.312026
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios · AAAI 2026
Natural language and speech › Language models and text generation › natural language understanding › question answering
spoken question answering
0.312026
End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering · AAAI 2026
Natural language and speech › Language models and text generation
in-context learning
0.312025
Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding · EMNLP 2025
Natural language and speech › Language models and text generation
long context
0.312025
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression · ACL (1) 2025
Machine learning › Efficient and distributed learning › inference efficiency
memory-efficient inference
0.312025
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 2.0contrastive language-speech pretraining · 2.0speculative decoding · 1.0non-autoregressive generation · 1.0bidirectional attention · 1.0rotary position embedding · 0.9layer-wise cache balancing · 0.9latent space down-sampling · 0.9entropy-based compression · 0.9attention-aware compression · 0.9
YearPublicationVenuePosition
2026 End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering
abstract
Significant progress has been made in spoken question answering (SQA) in recent years. However, many existing methods, including large audio language models, struggle with processing long audio. Follow the success of retrieval augmented generation, a speech-related retriever shows promising in help preprocessing long-form speech. But the performance of existing speech-related retrievers is lacking. To address this challenge, we propose CLSR, an end-to-end contrastive language-speech retriever that efficiently extracts question-relevant segments from long audio recordings for downstream SQA task. Unlike conventional speech-text contrastive models, CLSR incorporates an intermediate step that converts acoustic features into text-like representations prior to alignment, thereby more effectively bridging the gap between modalities. Experimental results across four cross-modal retrieval datasets demonstrate that CLSR surpasses both end-to-end speech related retrievers and pipeline approaches combining speech recognition with text retrieval, providing a robust foundation for advancing practical long-form SQA applications.
Jiliang Hu 0001, Zuchao Li, Baoyuan Qi, Guoming Liu, Ping Wang 0028
AAAI4
2026 Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
abstract
Speculative decoding accelerates LLM inference by utilizing otherwise idle computational resources during memory-to-chip data transfer. Current speculative decoding methods typically assume a considerable amount of available computing power, then generate a complex and massive draft tree using a small autoregressive language model to improve overall prediction accuracy. However, methods like batching have been widely applied in mainstream model inference systems as a superior alternative to speculative decoding, as they compress the available idle computing power. Therefore, performing speculative decoding with low verification resources and low scheduling costs has become an important research problem. We believe that more capable models that allow for parallel generation on draft sequences are what we truly need. Recognizing the fundamental nature of draft models to only generate sequences of limited length, we propose SpecFormer, a novel architecture that integrates unidirectional and bidirectional attention mechanisms. SpecFormer combines the autoregressive model’s ability to extract information from the entire input sequence with the parallel generation benefits of non-autoregressive models. This design eliminates the reliance on large prefix trees and achieves consistent acceleration, even in large-batch scenarios. Through lossless speculative decoding experiments across models of various scales, we demonstrate that SpecFormer sets a new standard for scaling LLM inference with lower training demands and reduced computational costs.
Luohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi, Guoming Liu, Hai Zhao 0001
AAAI5
2025 KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
abstract
Large language models (LLMs) based on Transformer Decoders have become the preferred choice for conversational generative AI.Despite the overall superiority of the Decoder architecture, the gradually increasing Key-Value (KV) cache during inference has emerged as a primary efficiency bottleneck, both in aspects of memory consumption and data transfer bandwidth limitations.To address these challenges, we propose a paradigm called KV-Latent.By down-sampling the Key-Value vector dimensions into a latent space, we can significantly reduce the KV Cache footprint and improve inference speed, only with a small amount of extra training, less than 1% of pretraining takes.Besides, we enhanced the stability of Rotary Positional Embedding applied on lower-dimensional vectors by modifying its frequency sampling mechanism, avoiding noise introduced by higher frequencies while retaining position attenuation.Our experiments, including both models with Grouped Query Attention and those without, have yielded satisfactory results.Finally, we conducted comparative experiments to study the impact of separately reducing Key and Value components on model's performance.Our approach allows for the construction of more efficient language model systems, and opens the new possibility on KV Cache saving and efficient LLMs.Our code is available at https://github.com/ShiLuohe/KV- Latent.
Luohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi, Guoming Liu, Hai Zhao 0001
ACL (1)5
2025 SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
abstract
Zicong Tang, Shi Luohe, Zuchao Li, Baoyuan Qi, Liu Guoming, Lefei Zhang, Ping Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zicong Tang, Luohe Shi, Zuchao Li, Baoyuan Qi, Guoming Liu, Lefei Zhang, Ping Wang 0028
ACL (1)5
2025 DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
abstract
Task-agnostic prompt compression leverages the redundancy in natural language to reduce computational overhead and enhance information density within prompts, especially in longcontext scenarios.Existing methods predominantly rely on information entropy as the metric to compress lexical units, aiming to achieve minimal information loss.However, these approaches overlook two critical aspects: (i) the importance of attention-critical tokens at the algorithmic level, and (ii) shifts in information entropy during the compression process.Motivated by these challenges, we propose a dynamic attention-aware approach for taskagnostic prompt compression (DAC).This approach effectively integrates entropy and attention information, dynamically sensing entropy shifts during compression to achieve fine-grained prompt compression.Extensive experiments across various domains, including LongBench, GSM8K, and BBH, show that DAC consistently yields robust and substantial improvements across a diverse range of tasks and LLMs, offering compelling evidence of its efficacy.
Zuchao Li, Hai Zhao 0001, Baoyuan Qi, Guoming Liu
ACL (1)5
2025 Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding
abstract
As a crucial method in prompt engineering, In-Context Learning (ICL) enhances the generalization and knowledge utilization capabilities of Large Language Models (LLMs) (Dong et al., 2024).However, the lengthy retrieved contexts and limited token throughput in autoregressive models significantly constrain reasoning speed.To address this challenge, we propose N-Gram Trie Speculative Decoding, a novel approach that leverages the overlap between context and model output.This method constructs an n-gram trie from the context to generate drafts, accelerating token generation for LLMs.We evaluate our approach on summarization, Retrieval-Augmented Generation (RAG), and contextbased Question Answering (QA) tasks.Experimental results on Vicuna-7B, Llama2-7B-Chat, and Llama3-8B-Instruct demonstrate substantial speed improvements without compromising accuracy.Compared with various strong baselines, our method achieves the highest mean speedup, showcasing its effectiveness and efficiency.Our implement code is available here: https://github.com/mrlife219/Ngram-Trie.
Jinglin Chen, Qiwei Li 0002, Zuchao Li, Baoyuan Qi, Guoming Liu, Haojun Ai, Hai Zhao 0001, Ping Wang 0028
EMNLP5
2025 XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks.However, their extensive memory requirements, particularly due to KV cache growth during long-text understanding and generation, present significant challenges for deployment in resourceconstrained environments.Quantization has emerged as a promising solution to reduce memory consumption while preserving historical information.We propose XQuant, a training-free and plug-and-play framework that achieves ultra-low equivalent bit-width KV cache quantization.XQuant introduces two key innovations: a computationally negligible data-free calibration method and cross-layer KV cache compression, enabling quantization to sub-1.4 bits.Extensive experiments on TruthfulQA and LongBench demonstrate that XQuant outperforms state-of-the-art methods (e.g., KIVI-2bit and AsymKV-1.5bit)by achieving lower bit-width while maintaining superior performance, establishing a better trade-off between memory efficiency and model accuracy.The source code is available at https: //github.com/brinenick511/XQuant.KeyCache[l][0], 14 KeyCache[l][1], 15 KeyCache[l][2] 16 else 17 DequantizedKey ← Dequantize 18 KeyCache[l -1][0], 19 KeyCache[l -1][1], 20 KeyCache[l][2] 21 if l < vm or l mod 2 == 0 then 22 DequantizedValue ← Dequantize 23 ValueCache[l][0], 24 ValueCache[l][1], 25 ValueCache[l][2] 26 else 27 DequantizedValue ← Dequantize 28 ValueCache[l -1][0], 29 ValueCache[l -1][1], 30 ValueCache[l][2]
Haoqi Yang 0001, Yao Yao 0008, Zuchao Li, Baoyuan Qi, Guoming Liu, Hai Zhao 0001
EMNLP5
2025 What Limits Bidirectional Model's Generative Capabilities? A Uni-Bi-Directional Mixture-of-Expert Method For Bidirectional Fine-tuning
abstract
Large Language Models (LLMs) excel in generation tasks, yet their causal attention mechanisms limit performance in embedding tasks. While bidirectional modeling may enhance embeddings, naively fine-tuning unidirectional models bidirectionally severely degrades generative performance. To investigate this trade-off, we analyze attention weights as dependence indicators and find that bidirectional fine-tuning increases subsequent dependence, impairing unidirectional generation. Through systematic Transformer module evaluations, we discover the FFN layer is least affected by such dependence. Leveraging this discovery, we propose UBMoE-LLM, a novel Uni-Bi-directional Mixture-of-Experts LLM, which integrates the original unidirectional FFN with a bidirectionally fine-tuned FFN via unsupervised contrastive learning. This MoE-based approach enhances embedding performance while preserving robust generation. Extensive experiments across diverse datasets and model scales validate our attention dependence metric and demonstrate UBMoE-LLM’s superior generative quality and reduced hallucination. Code is available at: https://github.com/heiyonghua/ubmoe_llm.
Zuchao Li, Yonghua Hei, Qiwei Li 0002, Lefei Zhang, Ping Wang 0028, Hai Zhao 0001, Baoyuan Qi, Guoming Liu
ICML8
2024 Cross-domain NER in the data-poor scenarios for human mobility knowledge
Fusheng Jin, Mengnan Chen, Guoming Liu, He Pang, Ye Yuan 0001
GeoInformatica4