Guangxuan Xiao

dblp:283/5633 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-7182-9284ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Efficient and distributed learning · 37% Deep learning architectures and training · 23% Language models and text generation · 17%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 100%

Topics — the 30 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
2.232024
BitDelta: Your Fine-Tune May Only Be Worth One Bit · NeurIPS 2024
Efficient Streaming Language Models with Attention Sinks · ICLR 2024
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models · ICML 2023
Machine learning › Deep learning architectures and training
attention mechanism
1.622025
Retrieval Head Mechanistically Explains Long-Context Factuality · ICLR 2025
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory · NeurIPS 2024
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression
1.622025
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads · ICLR 2025
QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference · ICML 2024
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model
1.622025
Retrieval Head Mechanistically Explains Long-Context Factuality · ICLR 2025
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model inference
long-context inference
1.122025
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads · ICLR 2025
QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference · ICML 2024
Machine learning › Efficient and distributed learning
inference efficiency
1.022024
QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference · ICML 2024
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory · NeurIPS 2024
Machine learning › Generative modeling
diffusion model
0.912025
FastComposer: Tuning-Free Multi-subject Image Generation with Localized Attention · Int. J. Comput. Vis. 2025
Machine learning › Deep learning architectures and training › transformer
efficient transformer
0.912025
XAttention: Block Sparse Attention with Antidiagonal Scoring · ICML 2025
Machine learning › Efficient and distributed learning › model compression › transformer compression › attention compression
KV cache pruning
0.912025
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads · ICLR 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling
0.912025
XAttention: Block Sparse Attention with Antidiagonal Scoring · ICML 2025
Machine learning › Generative modeling › diffusion model › controllable generation
multi-subject generation
0.912025
FastComposer: Tuning-Free Multi-subject Image Generation with Localized Attention · Int. J. Comput. Vis. 2025
Machine learning › Generative modeling › diffusion model › text-to-image generation
personalized text-to-image generation
0.912025
FastComposer: Tuning-Free Multi-subject Image Generation with Localized Attention · Int. J. Comput. Vis. 2025
Machine learning › Trustworthy machine learning › language model interpretability
retrieval heads
0.912025
Retrieval Head Mechanistically Explains Long-Context Factuality · ICLR 2025
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention
0.912025
XAttention: Block Sparse Attention with Antidiagonal Scoring · ICML 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
XAttention: Block Sparse Attention with Antidiagonal Scoring · ICML 2025
Machine learning › Trustworthy machine learning › language model interpretability
attention sink
0.812024
Efficient Streaming Language Models with Attention Sinks · ICLR 2024
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention
0.812024
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory · NeurIPS 2024
Machine learning › Efficient and distributed learning
embedding cache
0.812024
FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network Training · Proc. VLDB Endow. 2024
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.812024
BitDelta: Your Fine-Tune May Only Be Worth One Bit · NeurIPS 2024
Machine learning › Graph learning
graph neural network training
0.812024
FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network Training · Proc. VLDB Endow. 2024
Machine learning › Efficient and distributed learning
KV cache
0.812024
Efficient Streaming Language Models with Attention Sinks · ICLR 2024
Natural language and speech › Language models and text generation
long context
0.812024
Efficient Streaming Language Models with Attention Sinks · ICLR 2024
Machine learning › Efficient and distributed learning
memory-efficient training
0.812024
FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network Training · Proc. VLDB Endow. 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.812024
BitDelta: Your Fine-Tune May Only Be Worth One Bit · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression
quantization
0.812024
BitDelta: Your Fine-Tune May Only Be Worth One Bit · NeurIPS 2024
Machine learning › Deep learning architectures and training › attention mechanism › local attention
window attention
0.812024
Efficient Streaming Language Models with Attention Sinks · ICLR 2024
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization
0.712023
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models · ICML 2023
Machine learning › Deep learning architectures and training › attention mechanism › multi-head attention
attention head specialization
0.312025
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads · ICLR 2025
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.312025
Retrieval Head Mechanistically Explains Long-Context Factuality · ICLR 2025
Machine learning › Trustworthy machine learning
hallucination
0.312025
Retrieval Head Mechanistically Explains Long-Context Factuality · ICLR 2025

Methods — techniques the papers use, named apart from their topics

synthetic data · 0.9subject embedding · 0.9pruning · 0.9optimization-based head identification · 0.9delayed conditioning · 0.9cross-attention localization · 0.9block-sparse attention · 0.9attention head analysis · 0.9antidiagonal scoring · 0.9staleness criterion · 0.8gradient-based cache policy · 0.8attention sinks · 0.8
YearPublicationVenuePosition
2025 Retrieval Head Mechanistically Explains Long-Context Factuality
abstract
Despite the recent progress in long-context language models, it remains elusive how transformer-based models exhibit the capability to retrieve relevant information from arbitrary locations within the long context. This paper aims to address this question. Our systematic investigation across a wide spectrum of models reveals that a special type of attention heads are largely responsible for retrieving information, which we dub retrieval heads. We identify intriguing properties of retrieval heads:(1) universal: all the explored models with long-context capability have a set of retrieval heads; (2) sparse: only a small portion (less than 5\%) of the attention heads are retrieval. (3) intrinsic: retrieval heads already exist in models pretrained with short context. When extending the context length by continual pretraining, it is still the same set of heads that perform information retrieval. (4) dynamically activated: take Llama-2 7B for example, 12 retrieval heads always attend to the required information no matter how the context is changed. The rest of the retrieval heads are activated in different contexts. (5) causal: completely pruning retrieval heads leads to failure in retrieving relevant information and results in hallucination, while pruning random non-retrieval heads does not affect the model's retrieval ability. We further show that retrieval heads strongly influence chain-of-thought (CoT) reasoning, where the model needs to frequently refer back the question and previously-generated context. Conversely, tasks where the model directly generates the answer using its intrinsic knowledge are less impacted by masking out retrieval heads. These observations collectively explain which internal part of the model seeks information from the input tokens. We believe our insights will foster future research on reducing hallucination, improving reasoning, and compressing the KV cache.
Yizhong Wang, Guangxuan Xiao, Hao Peng 0018
ICLR3
2025 DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
abstract
Deploying long-context large language models (LLMs) is essential but poses significant computational and memory challenges. Caching all Key and Value (KV) states across all attention heads consumes substantial memory. Existing KV cache pruning methods either damage the long-context capabilities of LLMs or offer only limited efficiency improvements. In this paper, we identify that only a fraction of attention heads, a.k.a, Retrieval Heads, are critical for processing long contexts and require full attention across all tokens. In contrast, all other heads, which primarily focus on recent tokens and attention sinks—referred to as Streaming Heads—do not require full attention. Based on this insight, we introduce DuoAttention, a framework that only applies a full KV cache to retrieval heads while using a light-weight, constant-length KV cache for streaming heads, which reduces both LLM's decoding and pre-filling memory and latency without compromising its long-context abilities. DuoAttention uses a lightweight, optimization-based algorithm with synthetic data to identify retrieval heads accurately. Our method significantly reduces long-context inference memory by up to 2.55$\times$ for MHA and 1.67$\times$ for GQA models while speeding up decoding by up to 2.18$\times$ and 1.50$\times$ and accelerating pre-filling by up to 1.73$\times$ and 1.63$\times$ for MHA and GQA models, respectively, with minimal accuracy loss compared to full attention. Notably, combined with quantization, DuoAttention enables Llama-3-8B decoding with 3.33 million context length measured on a single A100 GPU. Code is provided in https://github.com/mit-han-lab/duo-attention.
Guangxuan Xiao, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Song Han 0003
ICLR1
2025 XAttention: Block Sparse Attention with Antidiagonal Scoring
abstract
Long-Context Transformer Models (LCTMs) are vital for real-world applications but suffer high computational costs due to attention’s quadratic complexity. Block-sparse attention mitigates this by focusing computation on critical regions, yet existing methods struggle with balancing accuracy and efficiency due to costly block importance measurements. In this paper, we introduce XAttention, a plug-and-play framework that dramatically accelerates long-context inference in Transformers models using sparse attention. XAttention’s key innovation is the insight that the sum of antidiagonal values (i.e., from the lower-left to upper-right) in the attention matrix provides a powerful proxy for block importance. This allows for precise identification and pruning of non-essential blocks, resulting in high sparsity and dramatically accelerated inference. Across comprehensive evaluations on demanding long-context benchmarks—including RULER and LongBench for language, VideoMME for video understanding, and VBench for video generation—XAttention achieves accuracy comparable to full attention while delivering substantial computational gains. We demonstrate up to 13.5x acceleration in attention computation. These results underscore XAttention’s ability to unlock the practical potential of block sparse attention, paving the way for scalable and efficient deployment of LCTMs in real-world applications.
Ruyi Xu, Guangxuan Xiao, Haofeng Huang, Junxian Guo, Song Han 0003
ICML2
2025 FastComposer: Tuning-Free Multi-subject Image Generation with Localized Attention
abstract
Abstract Diffusion models excel at text-to-image generation, especially in subject-driven generation for personalized images. However, existing methods are inefficient due to the subject-specific fine-tuning, which is computationally intensive and hampers efficient deployment. Moreover, existing methods struggle with multi-subject generation as they often blend identity among subjects. We present FastComposer which enables efficient, personalized, multi-subject text-to-image generation without fine-tuning. FastComposer uses subject embeddings extracted by an image encoder to augment the generic text conditioning in diffusion models, enabling personalized image generation based on subject images and textual instructions with only forward passes. To address the identity blending problem in the multi-subject generation, FastComposer proposes cross-attention localization supervision during training, enforcing the attention of reference subjects localized to the correct regions in the target images. Naively conditioning on subject embeddings results in subject overfitting. FastComposer proposes delayed subject conditioning in the denoising step to maintain both identity and editability in subject-driven image generation. FastComposer generates images of multiple unseen individuals with different styles, actions, and contexts. It achieves 300 $$\times $$ × –2500 $$\times $$ × speedup compared to fine-tuning-based methods and requires zero extra storage for new subjects. FastComposer paves the way for efficient, personalized, and high-quality multi-subject image creation. Code, model, and dataset are available here ( https://github.com/mit-han-lab/fastcomposer ).
Guangxuan Xiao, Tianwei Yin, William T. Freeman, Frédo Durand, Song Han 0003
Int. J. Comput. Vis.1
2024 Efficient Streaming Language Models with Attention Sinks
abstract
Deploying Large Language Models (LLMs) in streaming applications such as multi-round dialogue, where long interactions are expected, is urgently needed but poses two major challenges. Firstly, during the decoding stage, caching previous tokens' Key and Value states (KV) consumes extensive memory. Secondly, popular LLMs cannot generalize to longer texts than the training sequence length. Window attention, where only the most recent KVs are cached, is a natural approach --- but we show that it fails when the text length surpasses the cache size. We observe an interesting phenomenon, namely attention sink, that keeping the KV of initial tokens will largely recover the performance of window attention. In this paper, we first demonstrate that the emergence of attention sink is due to the strong attention scores towards initial tokens as a ``sink'' even if they are not semantically important. Based on the above analysis, we introduce StreamingLLM, an efficient framework that enables LLMs trained with a finite length attention window to generalize to infinite sequence length without any fine-tuning. We show that StreamingLLM can enable Llama-2, MPT, Falcon, and Pythia to perform stable and efficient language modeling with up to 4 million tokens and more. In addition, we discover that adding a placeholder token as a dedicated attention sink during pre-training can further improve streaming deployment. In streaming settings, StreamingLLM outperforms the sliding window recomputation baseline by up to 22.2$\times$ speedup. Code and datasets are provided at https://github.com/mit-han-lab/streaming-llm.
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 0003, Mike Lewis
ICLR1
2024 QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference
abstract
As the demand for long-context large language models (LLMs) increases, models with context windows of up to 128K or 1M tokens are becoming increasingly prevalent. However, long-context LLM inference is challenging since the inference speed decreases significantly as the sequence length grows. This slowdown is primarily caused by loading a large KV cache during self-attention. Previous works have shown that a small portion of critical tokens will dominate the attention outcomes. However, we observe the criticality of a token highly depends on the query. To this end, we propose Quest, a query-aware KV cache selection algorithm. Quest keeps track of the minimal and maximal Key values in KV cache pages and estimates the criticality of a given page using Query vectors. By only loading the Top-K critical KV cache pages for attention, Quest significantly speeds up self-attention without sacrificing accuracy. We show that Quest can achieve up to 2.23x self-attention speedup, which reduces inference latency by 7.03x while performing well on tasks with long dependencies with negligible accuracy loss. Code is available at https://github.com/mit-han-lab/quest.
Yilong Zhao 0002, Kan Zhu, Guangxuan Xiao, Baris Kasikci, Song Han 0003
ICML4
2024 BitDelta: Your Fine-Tune May Only Be Worth One Bit
abstract
Large Language Models (LLMs) are typically trained in two phases: pre-training on large internet-scale datasets, and fine-tuning for downstream tasks. Given the higher computational demand of pre-training, it is intuitive to assume that fine-tuning adds less new information to the model, and is thus more compressible. We explore this assumption by decomposing the weights of fine-tuned models into their pre-trained components and an additional delta. We introduce a simple method, BitDelta, which successfully quantizes this delta down to 1 bit without compromising performance. This interesting finding not only highlights the potential redundancy of information added during fine-tuning, but also has significant implications for the multi-tenant serving and multi-tenant storage of fine-tuned models. By enabling the use of a single high-precision base model accompanied by multiple 1-bit deltas, BitDelta dramatically reduces GPU memory requirements by more than 10x, thus reducing per-user generation latency by more than 10x in multi-tenant settings. We validate BitDelta through experiments across Llama-2, Mistral and MPT model families, and on models up to 70B parameters, showcasing minimal performance degradation in all tested settings.
James Liu, Guangxuan Xiao, Jason D. Lee, Song Han 0003, Tri Dao, Tianle Cai
NeurIPS2
2024 InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
abstract
Large language models (LLMs) have emerged as a cornerstone in real-world applications with lengthy streaming inputs (e.g., LLM-driven agents). However, existing LLMs, pre-trained on sequences with a restricted maximum length, cannot process longer sequences due to the out-of-domain and distraction issues. Common solutions often involve continual pre-training on longer sequences, which will introduce expensive computational overhead and uncontrollable change in model capabilities. In this paper, we unveil the intrinsic capacity of LLMs for understanding extremely long sequences without any fine-tuning. To this end, we introduce a training-free memory-based method, InfLLM. Specifically, InfLLM stores distant contexts into additional memory units and employs an efficient mechanism to lookup token-relevant units for attention computation. Thereby, InfLLM allows LLMs to efficiently process long sequences with a limited context window and well capture long-distance dependencies. Without any training, InfLLM enables LLMs that are pre-trained on sequences consisting of a few thousand tokens to achieve comparable performance with competitive baselines that continually train these LLMs on long sequences. Even when the sequence length is scaled to 1,024K, InfLLM still effectively captures long-distance dependencies. Our code can be found at https://github.com/thunlp/InfLLM.
Chaojun Xiao, Pengle Zhang, Xu Han 0007, Guangxuan Xiao, Yankai Lin 0001, Zhengyan Zhang, Zhiyuan Liu 0001, Maosong Sun 0001
NeurIPS4
2024 FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network Training
abstract
A key performance bottleneck when training graph neural network (GNN) models on large, real-world graphs is loading node features onto a GPU. Due to limited GPU memory, expensive data movement is necessary to facilitate the storage of these features on alternative devices with slower access (e.g. CPU memory). Moreover, the irregularity of graph structures contributes to poor data locality which further exacerbates the problem. Consequently, existing frameworks capable of efficiently training large GNN models usually incur a significant accuracy degradation because of the currently-available shortcuts involved. To address these limitations, we instead propose FreshGNN, a general-purpose GNN mini-batch training framework that leverages a historical cache for storing and reusing GNN node embeddings instead of re-computing them through fetching raw features at every iteration. Critical to its success, the corresponding cache policy is designed, using a combination of gradient-based and staleness criteria, to selectively screen those embeddings which are relatively stable and can be cached, from those that need to be re-computed to reduce estimation errors and subsequent downstream accuracy loss. When paired with complementary system enhancements to support this selective historical cache, FreshGNN is able to accelerate the training speed on large graph datasets such as ogbn-papers100M and MAG240M by 3.4× up to 20.5× and reduce the memory access by 59%, with less than 1% influence on test accuracy.
Kezhao Huang, Haitian Jiang, Guangxuan Xiao, David P. Wipf, Xiang Song 0003, Zengfeng Huang, Jidong Zhai, Zheng Zhang 0001
Proc. VLDB Endow.4
2023 SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
abstract
Large language models (LLMs) show excellent performance but are compute- and memory-intensive. Quantization can reduce memory and accelerate inference. However, existing methods cannot maintain accuracy and hardware efficiency at the same time. We propose SmoothQuant, a training-free, accuracy-preserving, and general-purpose post-training quantization (PTQ) solution to enable 8-bit weight, 8-bit activation (W8A8) quantization for LLMs. Based on the fact that weights are easy to quantize while activations are not, SmoothQuant smooths the activation outliers by offline migrating the quantization difficulty from activations to weights with a mathematically equivalent transformation. SmoothQuant enables an INT8 quantization of both weights and activations for all the matrix multiplications in LLMs, including OPT, BLOOM, GLM, MT-NLG, and LLaMA family. We demonstrate up to 1.56$\times$ speedup and 2$\times$ memory reduction for LLMs with negligible loss in accuracy. SmoothQuant enables serving 530B LLM within a single node. Our work offers a turn-key solution that reduces hardware costs and democratizes LLMs.
Guangxuan Xiao, Ji Lin 0002, Mickaël Seznec, Julien Demouth, Song Han 0003
ICML1