Baoyuan Qi

dblp:146/6949 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
11since 2021 · last 2026
0009-0008-5856-4817ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Efficient and distributed learning · 63% Language models and text generation · 32% Representation and self-supervised learning · 5%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference acceleration
1.922026
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios · AAAI 2026
Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding · EMNLP 2025
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding
1.922026
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios · AAAI 2026
Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding · EMNLP 2025
Machine learning › Efficient and distributed learning
inference efficiency
1.722025
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers · ACL (1) 2025
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding · ACL (1) 2025
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression
1.722025
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers · ACL (1) 2025
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding · ACL (1) 2025
Information retrieval
cross-modal retrieval
1.012026
End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering · AAAI 2026
Natural language and speech › Language models and text generation › neural language model
bidirectional language model
0.912025
What Limits Bidirectional Model's Generative Capabilities? A Uni-Bi-Directional Mixture-of-Expert Method For Bidirectional Fine-tuning · ICML 2025
Machine learning › Efficient and distributed learning
KV cache management
0.912025
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers · ACL (1) 2025
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization
0.912025
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025
Natural language and speech › Language models and text generation › large language model inference
long-context inference
0.912025
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025
Natural language and speech › Language models and text generation › language modeling
long-context language modeling
0.912025
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding · ACL (1) 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling
0.912025
SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers · ACL (1) 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025
Natural language and speech › Language models and text generation › prompting
prompt compression
0.912025
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression · ACL (1) 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.912025
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025
Machine learning › Representation and self-supervised learning
text embedding
0.912025
What Limits Bidirectional Model's Generative Capabilities? A Uni-Bi-Directional Mixture-of-Expert Method For Bidirectional Fine-tuning · ICML 2025
Natural language and speech › Language models and text generation
large language model inference
0.312026
Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios · AAAI 2026
Natural language and speech › Language models and text generation › natural language understanding › question answering
spoken question answering
0.312026
End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering · AAAI 2026
Natural language and speech › Language models and text generation
in-context learning
0.312025
Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding · EMNLP 2025
Natural language and speech › Language models and text generation
long context
0.312025
DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression · ACL (1) 2025
Machine learning › Efficient and distributed learning › inference efficiency
memory-efficient inference
0.312025
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 2.0contrastive language-speech pretraining · 2.0speculative decoding · 1.0non-autoregressive generation · 1.0bidirectional attention · 1.0rotary position embedding · 0.9layer-wise cache balancing · 0.9latent space down-sampling · 0.9entropy-based compression · 0.9attention-aware compression · 0.9
YearPublicationVenuePosition
2026 End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering
abstract
Significant progress has been made in spoken question answering (SQA) in recent years. However, many existing methods, including large audio language models, struggle with processing long audio. Follow the success of retrieval augmented generation, a speech-related retriever shows promising in help preprocessing long-form speech. But the performance of existing speech-related retrievers is lacking. To address this challenge, we propose CLSR, an end-to-end contrastive language-speech retriever that efficiently extracts question-relevant segments from long audio recordings for downstream SQA task. Unlike conventional speech-text contrastive models, CLSR incorporates an intermediate step that converts acoustic features into text-like representations prior to alignment, thereby more effectively bridging the gap between modalities. Experimental results across four cross-modal retrieval datasets demonstrate that CLSR surpasses both end-to-end speech related retrievers and pipeline approaches combining speech recognition with text retrieval, providing a robust foundation for advancing practical long-form SQA applications.
Jiliang Hu 0001, Zuchao Li, Baoyuan Qi, Guoming Liu, Ping Wang 0028
AAAI3
2026 Scaling LLM Speculative Decoding: Non-Autoregressive Forecasting in Large-Batch Scenarios
abstract
Speculative decoding accelerates LLM inference by utilizing otherwise idle computational resources during memory-to-chip data transfer. Current speculative decoding methods typically assume a considerable amount of available computing power, then generate a complex and massive draft tree using a small autoregressive language model to improve overall prediction accuracy. However, methods like batching have been widely applied in mainstream model inference systems as a superior alternative to speculative decoding, as they compress the available idle computing power. Therefore, performing speculative decoding with low verification resources and low scheduling costs has become an important research problem. We believe that more capable models that allow for parallel generation on draft sequences are what we truly need. Recognizing the fundamental nature of draft models to only generate sequences of limited length, we propose SpecFormer, a novel architecture that integrates unidirectional and bidirectional attention mechanisms. SpecFormer combines the autoregressive model’s ability to extract information from the entire input sequence with the parallel generation benefits of non-autoregressive models. This design eliminates the reliance on large prefix trees and achieves consistent acceleration, even in large-batch scenarios. Through lossless speculative decoding experiments across models of various scales, we demonstrate that SpecFormer sets a new standard for scaling LLM inference with lower training demands and reduced computational costs.
Luohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi, Guoming Liu, Hai Zhao 0001
AAAI4
2026 Denoising-Enhanced Multimodal Detection Network for Fake News Video Detection
abstract
The proliferation of fake news on short video platforms has recently attracted widespread attention, prompting a growing interest in fake news video detection. While existing studies show initial success, they may be constrained by the following limitations: (1) Current detection approaches lack logical understanding and integration capabilities, thereby leading to models missing implicit essential news information. (2) They simultaneously exploit news content for detection but fail to effectively distinguish between informative and non-informative elements, causing the attention bias within the detection model. To address these limitations, we propose a Denoising-Enhanced Multimodal Detection Network (D2et) to mine implicit essential news information and filter out redundant information. Firstly, to improve comprehension of essential content, we leverage the large language model to mine tags from two levels: (1) Global-Insight Tag, which captures the overall content of the short video; and (2) Local-Veracity Tag, which identifies essential details. We obtain the global and local level tags to jointly extract the explicit and structured content for clarifying essential information. Secondly, we design an information denoising mechanism to eliminate redundant noise. This mechanism refines multimodal features guided by the extracted tags, ensuring that the model focuses on relevant content and suppresses redundant information, thereby improving detection performance. Extensive experiments on two datasets demonstrate the effectiveness of our method, validating the contributions of multi-level tags and the information denoising mechanism.
Xueshan Deng, Penghang Yu, Shengsheng Qian, Baoyuan Qi, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.4
2026 Social-Aware Multimodal Graph Learning for Micro-Video Popularity Prediction
abstract
Micro-videos, characterized by their brief duration and rich multi-modal content, have become increasingly popular on digital platforms. Micro-Video Popularity Prediction (MVPP) is crucial for understanding user behavior, identifying emerging trends, and leveraging commercial opportunities. While existing MVPP methods predominantly focus on extracting features from the multi-modal content of micro-videos, they often overlook the valuable social interactions that occur within micro-video platforms. In this paper, we introduce a novel Social-Aware Multimodal Graph Learning (SAMGL) approach that integrates multi-modal content and social information to enhance MVPP. First, we propose Social-Aware Multimodal Precomputation (SAMP) to efficiently utilize complete graph information in micro-video popularity prediction, aggregating multi-modal signals to mitigate information loss. Second, we introduce Retrieval Augmentation Prediction (RAP), which retrieves similar micro-videos to enhance prediction by incorporating global content features beyond direct graph connections. To facilitate our approach, we construct the datasets by incorporating social interactions as graph information. Extensive experiments on the datasets demonstrate that our model achieves state-of-the-art performance, validating the effectiveness of integrating social information and the proposed local-global aggregation mechanisms.
Shangheng Chen, Shengsheng Qian, Jun Hu 0016, Quan Fang, Baoyuan Qi, Changsheng Xu
IEEE Trans. Multim.5
2026 Code-Driven LLM Agent for One-Shot Explanatory Visual Question Answering
abstract
Code-driven Large Language Models (LLMs) integrate both natural and formal languages, enhancing reasoning, precision, and interaction with execution environments, which in turn augments the capabilities of intelligent agents. Recent advancements in code-driven LLMs have proven pivotal for visual tasks, such as Visual Question Answering (VQA), a critical task at the intersection of computer vision and natural language processing. Despite significant progress, interpretability remains a challenge for VQA models, leading to the emergence of multimodal explanations for VQA. In this article, we propose the One-Shot and Training-Free Code-Driven LLM Agent (OneCoLA), a novel framework for Multimodal Explanatory Visual Question Answering (MEVQA). OneCoLA enables LLMs to generate multimodal explanations for the VQA task by utilizing a one-shot prompt to convert input questions into Python programs that model the reasoning process. The framework supports the flexible integration of open-world tools, ensuring adaptability to different problem contexts. During program execution, OneCoLA captures and preserves key execution data to enhance the interpretability of the results. Additionally, through further one-shot prompting, the framework generates multimodal explanations by combining execution outcomes with relevant visual content, providing both textual and visual context. Experimental results compared with state-of-the-art methods demonstrate the effectiveness of OneCoLA in generating accurate and interpretable multimodal explanations without the need for extensive training.
Zuyi Zhou, Dizhan Xue, Baoyuan Qi, Shengsheng Qian, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.3
2025 KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding
abstract
Large language models (LLMs) based on Transformer Decoders have become the preferred choice for conversational generative AI.Despite the overall superiority of the Decoder architecture, the gradually increasing Key-Value (KV) cache during inference has emerged as a primary efficiency bottleneck, both in aspects of memory consumption and data transfer bandwidth limitations.To address these challenges, we propose a paradigm called KV-Latent.By down-sampling the Key-Value vector dimensions into a latent space, we can significantly reduce the KV Cache footprint and improve inference speed, only with a small amount of extra training, less than 1% of pretraining takes.Besides, we enhanced the stability of Rotary Positional Embedding applied on lower-dimensional vectors by modifying its frequency sampling mechanism, avoiding noise introduced by higher frequencies while retaining position attenuation.Our experiments, including both models with Grouped Query Attention and those without, have yielded satisfactory results.Finally, we conducted comparative experiments to study the impact of separately reducing Key and Value components on model's performance.Our approach allows for the construction of more efficient language model systems, and opens the new possibility on KV Cache saving and efficient LLMs.Our code is available at https://github.com/ShiLuohe/KV- Latent.
Luohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi, Guoming Liu, Hai Zhao 0001
ACL (1)4
2025 SpindleKV: A Novel KV Cache Reduction Method Balancing Both Shallow and Deep Layers
abstract
Zicong Tang, Shi Luohe, Zuchao Li, Baoyuan Qi, Liu Guoming, Lefei Zhang, Ping Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zicong Tang, Luohe Shi, Zuchao Li, Baoyuan Qi, Guoming Liu, Lefei Zhang, Ping Wang 0028
ACL (1)4
2025 DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression
abstract
Task-agnostic prompt compression leverages the redundancy in natural language to reduce computational overhead and enhance information density within prompts, especially in longcontext scenarios.Existing methods predominantly rely on information entropy as the metric to compress lexical units, aiming to achieve minimal information loss.However, these approaches overlook two critical aspects: (i) the importance of attention-critical tokens at the algorithmic level, and (ii) shifts in information entropy during the compression process.Motivated by these challenges, we propose a dynamic attention-aware approach for taskagnostic prompt compression (DAC).This approach effectively integrates entropy and attention information, dynamically sensing entropy shifts during compression to achieve fine-grained prompt compression.Extensive experiments across various domains, including LongBench, GSM8K, and BBH, show that DAC consistently yields robust and substantial improvements across a diverse range of tasks and LLMs, offering compelling evidence of its efficacy.
Zuchao Li, Hai Zhao 0001, Baoyuan Qi, Guoming Liu
ACL (1)4
2025 Faster In-Context Learning for LLMs via N-Gram Trie Speculative Decoding
abstract
As a crucial method in prompt engineering, In-Context Learning (ICL) enhances the generalization and knowledge utilization capabilities of Large Language Models (LLMs) (Dong et al., 2024).However, the lengthy retrieved contexts and limited token throughput in autoregressive models significantly constrain reasoning speed.To address this challenge, we propose N-Gram Trie Speculative Decoding, a novel approach that leverages the overlap between context and model output.This method constructs an n-gram trie from the context to generate drafts, accelerating token generation for LLMs.We evaluate our approach on summarization, Retrieval-Augmented Generation (RAG), and contextbased Question Answering (QA) tasks.Experimental results on Vicuna-7B, Llama2-7B-Chat, and Llama3-8B-Instruct demonstrate substantial speed improvements without compromising accuracy.Compared with various strong baselines, our method achieves the highest mean speedup, showcasing its effectiveness and efficiency.Our implement code is available here: https://github.com/mrlife219/Ngram-Trie.
Jinglin Chen, Qiwei Li 0002, Zuchao Li, Baoyuan Qi, Guoming Liu, Haojun Ai, Hai Zhao 0001, Ping Wang 0028
EMNLP4
2025 XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression
abstract
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse natural language processing tasks.However, their extensive memory requirements, particularly due to KV cache growth during long-text understanding and generation, present significant challenges for deployment in resourceconstrained environments.Quantization has emerged as a promising solution to reduce memory consumption while preserving historical information.We propose XQuant, a training-free and plug-and-play framework that achieves ultra-low equivalent bit-width KV cache quantization.XQuant introduces two key innovations: a computationally negligible data-free calibration method and cross-layer KV cache compression, enabling quantization to sub-1.4 bits.Extensive experiments on TruthfulQA and LongBench demonstrate that XQuant outperforms state-of-the-art methods (e.g., KIVI-2bit and AsymKV-1.5bit)by achieving lower bit-width while maintaining superior performance, establishing a better trade-off between memory efficiency and model accuracy.The source code is available at https: //github.com/brinenick511/XQuant.KeyCache[l][0], 14 KeyCache[l][1], 15 KeyCache[l][2] 16 else 17 DequantizedKey ← Dequantize 18 KeyCache[l -1][0], 19 KeyCache[l -1][1], 20 KeyCache[l][2] 21 if l < vm or l mod 2 == 0 then 22 DequantizedValue ← Dequantize 23 ValueCache[l][0], 24 ValueCache[l][1], 25 ValueCache[l][2] 26 else 27 DequantizedValue ← Dequantize 28 ValueCache[l -1][0], 29 ValueCache[l -1][1], 30 ValueCache[l][2]
Haoqi Yang 0001, Yao Yao 0008, Zuchao Li, Baoyuan Qi, Guoming Liu, Hai Zhao 0001
EMNLP4
2025 What Limits Bidirectional Model's Generative Capabilities? A Uni-Bi-Directional Mixture-of-Expert Method For Bidirectional Fine-tuning
abstract
Large Language Models (LLMs) excel in generation tasks, yet their causal attention mechanisms limit performance in embedding tasks. While bidirectional modeling may enhance embeddings, naively fine-tuning unidirectional models bidirectionally severely degrades generative performance. To investigate this trade-off, we analyze attention weights as dependence indicators and find that bidirectional fine-tuning increases subsequent dependence, impairing unidirectional generation. Through systematic Transformer module evaluations, we discover the FFN layer is least affected by such dependence. Leveraging this discovery, we propose UBMoE-LLM, a novel Uni-Bi-directional Mixture-of-Experts LLM, which integrates the original unidirectional FFN with a bidirectionally fine-tuned FFN via unsupervised contrastive learning. This MoE-based approach enhances embedding performance while preserving robust generation. Extensive experiments across diverse datasets and model scales validate our attention dependence metric and demonstrate UBMoE-LLM’s superior generative quality and reduced hallucination. Code is available at: https://github.com/heiyonghua/ubmoe_llm.
Zuchao Li, Yonghua Hei, Qiwei Li 0002, Lefei Zhang, Ping Wang 0028, Hai Zhao 0001, Baoyuan Qi, Guoming Liu
ICML7
2017 Attentive Path Combination for Knowledge Graph Completion
abstract
Knowledge graphs (KGs) are often significantly incomplete, necessitating a demand for KG completion. Path-based relation inference is one of the most important approaches to this task. Traditional methods treat each path between entity pairs as an atomic feature, thus inducing sparsity. Recently, neural network models solve this problem by decomposing a path as the sequence of relations in the path, before modelling path representations with Recurrent Neural Network (RNN) architectures. In cases there are multiple paths between an entity pair, state-of-the-art neural models either select only one path, or make usage of simple score pooling methods like Top-K, Average, LogSumExp. Unfortunately, none of these methods can model the scenario where relations can only be inferred by considering multiple informative paths collectively. In this paper, we propose a novel path-based relation inference model that learns entity pair representations with attentive path combination. Given an entity pair and a set of paths connecting the pair, our model allows for integrating information from each informative path, and form a dynamic entity pair representation for each query relation. We empirically evaluate the proposed method on a real-world dataset. Experimental results show that the proposed model achieves better performance than state-of-the-art path-based relation inference methods.
Quan Wang 0002, Baoyuan Qi, Yongqin Qiu, Peng Li 0021, Bin Wang 0004
ACML3
2015 An Environment Visual Awareness Approach in Cognitive Model ABGP
abstract
ABGP is a special cognitive model, which consists of awareness, beliefs, goals and plans. As most agent architectures, ABGP agents obtain knowledge from the natural scenes only through single preestablished rules as well, don't directly capture the natural scenes information like human visual. Inspired by the biological visual cortex (V1) and the higher brain areas perceiving visual features, we propose a novel deep network model convolutional generative stochastic model (CGSM) used to visual feature representation, and firstly introduce it into the awareness module of the cognitive model ABGP to construct a state-of-the-art cognitive model ABGP-CGSM. For the novel cognitive model ABGP-CGSM, we construct a rat-robot maze search simulation platform to show the validity recognizing natural scenes. According to the simulation results on the noise and noiseless natural scenes, the rat-robot implemented by ABGP-CGSM has an excellent success rate when passing through the maze. The simulation shows that the ABGP-CGSM model proposed in our work can directly enhance the capability of communication between agent and natural scenes, improve the ability to cognize the real world as human being and conduct the agent to plan independently its path in terms of the visual information from the natural scenes.
Gang Ma 0001, Bo Zhang 0003, Baoyuan Qi, Zhongzhi Shi
ICTAI4