EDBT 2026 Demo / reviewers in the wild / expert
Guangxuan Xiao
dblp:283/5633
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-7182-9284ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Efficient and distributed learning · 37% Deep learning architectures and training · 23% Language models and text generation · 17% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Memory systems · 100% |
Topics — the 30 heaviest of 33, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
2.2 | 3 | 2024 | BitDelta: Your Fine-Tune May Only Be Worth One Bit · NeurIPS 2024 Efficient Streaming Language Models with Attention Sinks · ICLR 2024 SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models · ICML 2023 |
Machine learning › Deep learning architectures and training
attention mechanism |
1.6 | 2 | 2025 | Retrieval Head Mechanistically Explains Long-Context Factuality · ICLR 2025 InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression |
1.6 | 2 | 2025 | DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads · ICLR 2025 QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference · ICML 2024 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model |
1.6 | 2 | 2025 | Retrieval Head Mechanistically Explains Long-Context Factuality · ICLR 2025 InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model inference
long-context inference |
1.1 | 2 | 2025 | DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads · ICLR 2025 QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference · ICML 2024 |
Machine learning › Efficient and distributed learning
inference efficiency |
1.0 | 2 | 2024 | QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference · ICML 2024 InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory · NeurIPS 2024 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | FastComposer: Tuning-Free Multi-subject Image Generation with Localized Attention · Int. J. Comput. Vis. 2025 |
Machine learning › Deep learning architectures and training › transformer
efficient transformer |
0.9 | 1 | 2025 | XAttention: Block Sparse Attention with Antidiagonal Scoring · ICML 2025 |
Machine learning › Efficient and distributed learning › model compression › transformer compression › attention compression
KV cache pruning |
0.9 | 1 | 2025 | DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads · ICLR 2025 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization
long-context modeling |
0.9 | 1 | 2025 | XAttention: Block Sparse Attention with Antidiagonal Scoring · ICML 2025 |
Machine learning › Generative modeling › diffusion model › controllable generation
multi-subject generation |
0.9 | 1 | 2025 | FastComposer: Tuning-Free Multi-subject Image Generation with Localized Attention · Int. J. Comput. Vis. 2025 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
personalized text-to-image generation |
0.9 | 1 | 2025 | FastComposer: Tuning-Free Multi-subject Image Generation with Localized Attention · Int. J. Comput. Vis. 2025 |
Machine learning › Trustworthy machine learning › language model interpretability
retrieval heads |
0.9 | 1 | 2025 | Retrieval Head Mechanistically Explains Long-Context Factuality · ICLR 2025 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
0.9 | 1 | 2025 | XAttention: Block Sparse Attention with Antidiagonal Scoring · ICML 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.9 | 1 | 2025 | XAttention: Block Sparse Attention with Antidiagonal Scoring · ICML 2025 |
Machine learning › Trustworthy machine learning › language model interpretability
attention sink |
0.8 | 1 | 2024 | Efficient Streaming Language Models with Attention Sinks · ICLR 2024 |
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention |
0.8 | 1 | 2024 | InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
embedding cache |
0.8 | 1 | 2024 | FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network Training · Proc. VLDB Endow. 2024 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.8 | 1 | 2024 | BitDelta: Your Fine-Tune May Only Be Worth One Bit · NeurIPS 2024 |
Machine learning › Graph learning
graph neural network training |
0.8 | 1 | 2024 | FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network Training · Proc. VLDB Endow. 2024 |
Machine learning › Efficient and distributed learning
KV cache |
0.8 | 1 | 2024 | Efficient Streaming Language Models with Attention Sinks · ICLR 2024 |
Natural language and speech › Language models and text generation
long context |
0.8 | 1 | 2024 | Efficient Streaming Language Models with Attention Sinks · ICLR 2024 |
Machine learning › Efficient and distributed learning
memory-efficient training |
0.8 | 1 | 2024 | FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network Training · Proc. VLDB Endow. 2024 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.8 | 1 | 2024 | BitDelta: Your Fine-Tune May Only Be Worth One Bit · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.8 | 1 | 2024 | BitDelta: Your Fine-Tune May Only Be Worth One Bit · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › attention mechanism › local attention
window attention |
0.8 | 1 | 2024 | Efficient Streaming Language Models with Attention Sinks · ICLR 2024 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.7 | 1 | 2023 | SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models · ICML 2023 |
Machine learning › Deep learning architectures and training › attention mechanism › multi-head attention
attention head specialization |
0.3 | 1 | 2025 | DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads · ICLR 2025 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.3 | 1 | 2025 | Retrieval Head Mechanistically Explains Long-Context Factuality · ICLR 2025 |
Machine learning › Trustworthy machine learning
hallucination |
0.3 | 1 | 2025 | Retrieval Head Mechanistically Explains Long-Context Factuality · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
synthetic data · 0.9subject embedding · 0.9pruning · 0.9optimization-based head identification · 0.9delayed conditioning · 0.9cross-attention localization · 0.9block-sparse attention · 0.9attention head analysis · 0.9antidiagonal scoring · 0.9staleness criterion · 0.8gradient-based cache policy · 0.8attention sinks · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Retrieval Head Mechanistically Explains Long-Context FactualityabstractDespite the recent progress in long-context language models, it remains elusive how transformer-based models exhibit the capability to retrieve relevant information from arbitrary locations within the long context. This paper aims to address this question. Our systematic investigation across a wide spectrum of models reveals that a special type of attention heads are largely responsible for retrieving information, which we dub retrieval heads. We identify intriguing properties of retrieval heads:(1) universal: all the explored models with long-context capability have a set of retrieval heads; (2) sparse: only a small portion (less than 5\%) of the attention heads are retrieval. (3) intrinsic: retrieval heads already exist in models pretrained with short context. When extending the context length by continual pretraining, it is still the same set of heads that perform information retrieval. (4) dynamically activated: take Llama-2 7B for example, 12 retrieval heads always attend to the required information no matter how the context is changed. The rest of the retrieval heads are activated in different contexts. (5) causal: completely pruning retrieval heads leads to failure in retrieving relevant information and results in hallucination, while pruning random non-retrieval heads does not affect the model's retrieval ability. We further show that retrieval heads strongly influence chain-of-thought (CoT) reasoning, where the model needs to frequently refer back the question and previously-generated context. Conversely, tasks where the model directly generates the answer using its intrinsic knowledge are less impacted by masking out retrieval heads. These observations collectively explain which internal part of the model seeks information from the input tokens. We believe our insights will foster future research on reducing hallucination, improving reasoning, and compressing the KV cache. Yizhong Wang, Guangxuan Xiao, Hao Peng 0018 |
ICLR | 3 |
| 2025 | DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming HeadsabstractDeploying long-context large language models (LLMs) is essential but poses significant computational and memory challenges.
Caching all Key and Value (KV) states across all attention heads consumes substantial memory.
Existing KV cache pruning methods either damage the long-context capabilities of LLMs or offer only limited efficiency improvements.
In this paper, we identify that only a fraction of attention heads, a.k.a, Retrieval Heads, are critical for processing long contexts and require full attention across all tokens.
In contrast, all other heads, which primarily focus on recent tokens and attention sinks—referred to as Streaming Heads—do not require full attention.
Based on this insight, we introduce DuoAttention, a framework that only applies a full KV cache to retrieval heads while using a light-weight, constant-length KV cache for streaming heads, which reduces both LLM's decoding and pre-filling memory and latency without compromising its long-context abilities.
DuoAttention uses a lightweight, optimization-based algorithm with synthetic data to identify retrieval heads accurately.
Our method significantly reduces long-context inference memory by up to 2.55$\times$ for MHA and 1.67$\times$ for GQA models while speeding up decoding by up to 2.18$\times$ and 1.50$\times$ and accelerating pre-filling by up to 1.73$\times$ and 1.63$\times$ for MHA and GQA models, respectively, with minimal accuracy loss compared to full attention.
Notably, combined with quantization, DuoAttention enables Llama-3-8B decoding with 3.33 million context length measured on a single A100 GPU. Code is provided in https://github.com/mit-han-lab/duo-attention. Guangxuan Xiao, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Song Han 0003 |
ICLR | 1 |
| 2025 | XAttention: Block Sparse Attention with Antidiagonal ScoringabstractLong-Context Transformer Models (LCTMs) are vital for real-world applications but suffer high computational costs due to attention’s quadratic complexity. Block-sparse attention mitigates this by focusing computation on critical regions, yet existing methods struggle with balancing accuracy and efficiency due to costly block importance measurements. In this paper, we introduce XAttention, a plug-and-play framework that dramatically accelerates long-context inference in Transformers models using sparse attention. XAttention’s key innovation is the insight that the sum of antidiagonal values (i.e., from the lower-left to upper-right) in the attention matrix provides a powerful proxy for block importance. This allows for precise identification and pruning of non-essential blocks, resulting in high sparsity and dramatically accelerated inference. Across comprehensive evaluations on demanding long-context benchmarks—including RULER and LongBench for language, VideoMME for video understanding, and VBench for video generation—XAttention achieves accuracy comparable to full attention while delivering substantial computational gains. We demonstrate up to 13.5x acceleration in attention computation. These results underscore XAttention’s ability to unlock the practical potential of block sparse attention, paving the way for scalable and efficient deployment of LCTMs in real-world applications. Ruyi Xu, Guangxuan Xiao, Haofeng Huang, Junxian Guo, Song Han 0003 |
ICML | 2 |
| 2025 | FastComposer: Tuning-Free Multi-subject Image Generation with Localized AttentionabstractAbstract Diffusion models excel at text-to-image generation, especially in subject-driven generation for personalized images. However, existing methods are inefficient due to the subject-specific fine-tuning, which is computationally intensive and hampers efficient deployment. Moreover, existing methods struggle with multi-subject generation as they often blend identity among subjects. We present FastComposer which enables efficient, personalized, multi-subject text-to-image generation without fine-tuning. FastComposer uses subject embeddings extracted by an image encoder to augment the generic text conditioning in diffusion models, enabling personalized image generation based on subject images and textual instructions with only forward passes. To address the identity blending problem in the multi-subject generation, FastComposer proposes cross-attention localization supervision during training, enforcing the attention of reference subjects localized to the correct regions in the target images. Naively conditioning on subject embeddings results in subject overfitting. FastComposer proposes delayed subject conditioning in the denoising step to maintain both identity and editability in subject-driven image generation. FastComposer generates images of multiple unseen individuals with different styles, actions, and contexts. It achieves 300 $$\times $$ × –2500 $$\times $$ × speedup compared to fine-tuning-based methods and requires zero extra storage for new subjects. FastComposer paves the way for efficient, personalized, and high-quality multi-subject image creation. Code, model, and dataset are available here ( https://github.com/mit-han-lab/fastcomposer ). Guangxuan Xiao, Tianwei Yin, William T. Freeman, Frédo Durand, Song Han 0003 |
Int. J. Comput. Vis. | 1 |
| 2024 | Efficient Streaming Language Models with Attention SinksabstractDeploying Large Language Models (LLMs) in streaming applications such as multi-round dialogue, where long interactions are expected, is urgently needed but poses two major challenges.
Firstly, during the decoding stage, caching previous tokens' Key and Value states (KV) consumes extensive memory.
Secondly, popular LLMs cannot generalize to longer texts than the training sequence length.
Window attention, where only the most recent KVs are cached, is a natural approach --- but we show that it fails when the text length surpasses the cache size.
We observe an interesting phenomenon, namely attention sink, that keeping the KV of initial tokens will largely recover the performance of window attention. In this paper, we first demonstrate that the emergence of attention sink is due to the strong attention scores towards initial tokens as a ``sink'' even if they are not semantically important.
Based on the above analysis, we introduce StreamingLLM, an efficient framework that enables LLMs trained with a finite length attention window to generalize to infinite sequence length without any fine-tuning.
We show that StreamingLLM can enable Llama-2, MPT, Falcon, and Pythia to perform stable and efficient language modeling with up to 4 million tokens and more.
In addition, we discover that adding a placeholder token as a dedicated attention sink during pre-training can further improve streaming deployment. In streaming settings, StreamingLLM outperforms the sliding window recomputation baseline by up to 22.2$\times$ speedup.
Code and datasets are provided at https://github.com/mit-han-lab/streaming-llm. Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 0003, Mike Lewis |
ICLR | 1 |
| 2024 | QUEST: Query-Aware Sparsity for Efficient Long-Context LLM InferenceabstractAs the demand for long-context large language models (LLMs) increases, models with context windows of up to 128K or 1M tokens are becoming increasingly prevalent. However, long-context LLM inference is challenging since the inference speed decreases significantly as the sequence length grows. This slowdown is primarily caused by loading a large KV cache during self-attention. Previous works have shown that a small portion of critical tokens will dominate the attention outcomes. However, we observe the criticality of a token highly depends on the query. To this end, we propose Quest, a query-aware KV cache selection algorithm. Quest keeps track of the minimal and maximal Key values in KV cache pages and estimates the criticality of a given page using Query vectors. By only loading the Top-K critical KV cache pages for attention, Quest significantly speeds up self-attention without sacrificing accuracy. We show that Quest can achieve up to 2.23x self-attention speedup, which reduces inference latency by 7.03x while performing well on tasks with long dependencies with negligible accuracy loss. Code is available at https://github.com/mit-han-lab/quest. Yilong Zhao 0002, Kan Zhu, Guangxuan Xiao, Baris Kasikci, Song Han 0003 |
ICML | 4 |
| 2024 | BitDelta: Your Fine-Tune May Only Be Worth One BitabstractLarge Language Models (LLMs) are typically trained in two phases: pre-training on large internet-scale datasets, and fine-tuning for downstream tasks. Given the higher computational demand of pre-training, it is intuitive to assume that fine-tuning adds less new information to the model, and is thus more compressible. We explore this assumption by decomposing the weights of fine-tuned models into their pre-trained components and an additional delta. We introduce a simple method, BitDelta, which successfully quantizes this delta down to 1 bit without compromising performance. This interesting finding not only highlights the potential redundancy of information added during fine-tuning, but also has significant implications for the multi-tenant serving and multi-tenant storage of fine-tuned models. By enabling the use of a single high-precision base model accompanied by multiple 1-bit deltas, BitDelta dramatically reduces GPU memory requirements by more than 10x, thus reducing per-user generation latency by more than 10x in multi-tenant settings. We validate BitDelta through experiments across Llama-2, Mistral and MPT model families, and on models up to 70B parameters, showcasing minimal performance degradation in all tested settings. James Liu, Guangxuan Xiao, Jason D. Lee, Song Han 0003, Tri Dao, Tianle Cai |
NeurIPS | 2 |
| 2024 | InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context MemoryabstractLarge language models (LLMs) have emerged as a cornerstone in real-world applications with lengthy streaming inputs (e.g., LLM-driven agents). However, existing LLMs, pre-trained on sequences with a restricted maximum length, cannot process longer sequences due to the out-of-domain and distraction issues. Common solutions often involve continual pre-training on longer sequences, which will introduce expensive computational overhead and uncontrollable change in model capabilities. In this paper, we unveil the intrinsic capacity of LLMs for understanding extremely long sequences without any fine-tuning. To this end, we introduce a training-free memory-based method, InfLLM. Specifically, InfLLM stores distant contexts into additional memory units and employs an efficient mechanism to lookup token-relevant units for attention computation. Thereby, InfLLM allows LLMs to efficiently process long sequences with a limited context window and well capture long-distance dependencies. Without any training, InfLLM enables LLMs that are pre-trained on sequences consisting of a few thousand tokens to achieve comparable performance with competitive baselines that continually train these LLMs on long sequences. Even when the sequence length is scaled to 1,024K, InfLLM still effectively captures long-distance dependencies. Our code can be found at https://github.com/thunlp/InfLLM. Chaojun Xiao, Pengle Zhang, Xu Han 0007, Guangxuan Xiao, Yankai Lin 0001, Zhengyan Zhang, Zhiyuan Liu 0001, Maosong Sun 0001 |
NeurIPS | 4 |
| 2024 | FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network TrainingabstractA key performance bottleneck when training graph neural network (GNN) models on large, real-world graphs is loading node features onto a GPU. Due to limited GPU memory, expensive data movement is necessary to facilitate the storage of these features on alternative devices with slower access (e.g. CPU memory). Moreover, the irregularity of graph structures contributes to poor data locality which further exacerbates the problem. Consequently, existing frameworks capable of efficiently training large GNN models usually incur a significant accuracy degradation because of the currently-available shortcuts involved. To address these limitations, we instead propose FreshGNN, a general-purpose GNN mini-batch training framework that leverages a historical cache for storing and reusing GNN node embeddings instead of re-computing them through fetching raw features at every iteration. Critical to its success, the corresponding cache policy is designed, using a combination of gradient-based and staleness criteria, to selectively screen those embeddings which are relatively stable and can be cached, from those that need to be re-computed to reduce estimation errors and subsequent downstream accuracy loss. When paired with complementary system enhancements to support this selective historical cache, FreshGNN is able to accelerate the training speed on large graph datasets such as ogbn-papers100M and MAG240M by 3.4× up to 20.5× and reduce the memory access by 59%, with less than 1% influence on test accuracy. Kezhao Huang, Haitian Jiang, Guangxuan Xiao, David P. Wipf, Xiang Song 0003, Zengfeng Huang, Jidong Zhai, Zheng Zhang 0001 |
Proc. VLDB Endow. | 4 |
| 2023 | SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsabstractLarge language models (LLMs) show excellent performance but are compute- and memory-intensive. Quantization can reduce memory and accelerate inference. However, existing methods cannot maintain accuracy and hardware efficiency at the same time. We propose SmoothQuant, a training-free, accuracy-preserving, and general-purpose post-training quantization (PTQ) solution to enable 8-bit weight, 8-bit activation (W8A8) quantization for LLMs. Based on the fact that weights are easy to quantize while activations are not, SmoothQuant smooths the activation outliers by offline migrating the quantization difficulty from activations to weights with a mathematically equivalent transformation. SmoothQuant enables an INT8 quantization of both weights and activations for all the matrix multiplications in LLMs, including OPT, BLOOM, GLM, MT-NLG, and LLaMA family. We demonstrate up to 1.56$\times$ speedup and 2$\times$ memory reduction for LLMs with negligible loss in accuracy. SmoothQuant enables serving 530B LLM within a single node. Our work offers a turn-key solution that reduces hardware costs and democratizes LLMs. Guangxuan Xiao, Ji Lin 0002, Mickaël Seznec, Julien Demouth, Song Han 0003 |
ICML | 1 |