VLDB 2026 Research / reviewers in the wild / expert
Coleman Hooper
dblp:279/6669 · also Coleman Richard Charles Hooper
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-5890-610XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Efficient and distributed learning · 57% Language models and text generation · 32% Deep learning architectures and training · 12% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 74% Memory systems · 26% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
model compression |
2.3 | 4 | 2025 | KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization · NeurIPS 2024 SqueezeLLM: Dense-and-Sparse Quantization · ICML 2024 EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference · MICRO 2021 |
Natural language and speech › Language models and text generation › large language model inference
long-context inference |
1.9 | 3 | 2025 | QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache · ICML 2025 KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization · NeurIPS 2024 Squeezed Attention: Accelerating Long Context Length LLM Inference · ACL (1) 2025 |
Machine learning › Efficient and distributed learning
inference efficiency |
1.7 | 2 | 2025 | QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache · ICML 2025 Squeezed Attention: Accelerating Long Context Length LLM Inference · ACL (1) 2025 |
Natural language and speech › Language models and text generation
large language model inference |
1.0 | 2 | 2024 | KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization · NeurIPS 2024 SqueezeLLM: Dense-and-Sparse Quantization · ICML 2024 |
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention |
0.9 | 1 | 2025 | Multipole Attention for Efficient Long Context Reasoning · NeurIPS 2025 |
Natural language and speech › Language models and text generation › large language model
large reasoning model |
0.9 | 1 | 2025 | Multipole Attention for Efficient Long Context Reasoning · NeurIPS 2025 |
Natural language and speech › Language models and text generation › chain-of-thought reasoning
long chain-of-thought reasoning |
0.9 | 1 | 2025 | Multipole Attention for Efficient Long Context Reasoning · NeurIPS 2025 |
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention |
0.9 | 1 | 2025 | Multipole Attention for Efficient Long Context Reasoning · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization |
0.8 | 1 | 2024 | KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression › quantization
non-uniform quantization |
0.8 | 1 | 2024 | SqueezeLLM: Dense-and-Sparse Quantization · ICML 2024 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.8 | 1 | 2024 | SqueezeLLM: Dense-and-Sparse Quantization · ICML 2024 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.8 | 1 | 2024 | KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression › sparsity
weight sparsity |
0.8 | 1 | 2024 | SqueezeLLM: Dense-and-Sparse Quantization · ICML 2024 |
Machine learning › Efficient and distributed learning › inference efficiency
energy-efficient inference |
0.5 | 1 | 2021 | EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference · MICRO 2021 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
transformer inference |
0.5 | 1 | 2021 | EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference · MICRO 2021 |
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression |
0.3 | 1 | 2025 | Multipole Attention for Efficient Long Context Reasoning · NeurIPS 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.2 | 1 | 2024 | KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization · NeurIPS 2024 |
Natural language and speech › Language models and text generation › pre-trained language model
BERT |
0.1 | 1 | 2021 | EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference · MICRO 2021 |
Methods — techniques the papers use, named apart from their topics
weight quantization · 1.7speculative decoding · 1.7KV cache quantization · 1.7kernel implementation · 0.9clustering · 0.9cluster centroid approximation · 0.9attention squeezing · 0.9sensitivity-based quantization · 0.8second-order information · 0.8per-channel quantization · 0.8non-uniform quantization · 0.8dense-and-sparse quantization · 0.8dense-and-sparse decomposition · 0.8custom CUDA kernels · 0.8sentence-level energy optimization · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Squeezed Attention: Accelerating Long Context Length LLM InferenceabstractColeman Richard Charles Hooper, Sehoon Kim, Hiva Mohammadzadeh, Monishwaran Maheswaran, Sebastian Zhao, June Paik, Michael W. Mahoney, Kurt Keutzer, Amir Gholami. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Coleman Hooper, Sehoon Kim 0001, Hiva Mohammadzadeh, Monishwaran Maheswaran, Sebastian Zhao, June Paik, Michael W. Mahoney, Kurt Keutzer, Amir Gholami |
ACL (1) | 1 |
| 2025 | QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV CacheabstractLarge Language Models (LLMs) are increasingly being deployed on edge devices for long-context settings, creating a growing need for fast and efficient long-context inference. In these scenarios, the Key-Value (KV) cache is the primary bottleneck in terms of both GPU memory and latency, as the full KV cache must be loaded for each decoding step. While speculative decoding is a widely accepted technique to accelerate autoregressive decoding, existing methods often struggle to achieve significant speedups due to inefficient KV cache optimization strategies and result in low acceptance rates. To address these challenges, we propose a novel self-speculative decoding framework, QuantSpec, where the draft model shares the architecture of the target model but employs a hierarchical 4-bit quantized KV cache and 4-bit quantized weights for acceleration. QuantSpec maintains high acceptance rates ($>$90\%) and reliably provides consistent end-to-end speedups upto $\sim2.5\times$, outperforming other self-speculative decoding methods that use sparse KV cache for long-context LLM inference. QuantSpec also reduces the memory requirements by $\sim 1.3\times$ compared to these alternatives. Rishabh Tiwari, Haocheng Xi, Aditya Tomar, Coleman Hooper, Sehoon Kim 0001, Maxwell Horton, Mahyar Najibi, Michael W. Mahoney, Kurt Keutzer, Amir Gholami |
ICML | 4 |
| 2025 | Multipole Attention for Efficient Long Context ReasoningabstractLarge Reasoning Models (LRMs) have shown promising accuracy improvements on complex problem-solving tasks. While these models have attained high accuracy by leveraging additional computation at test time, they need to generate long chain-of-thought reasoning in order to think before answering, which requires generating thousands of tokens.
While sparse attention methods can help reduce the KV cache pressure induced by this long autoregressive reasoning, these methods can introduce errors which disrupt the reasoning process.
Our work addresses these challenges by introducing Multipole Attention, which accelerates autoregressive reasoning by only computing exact attention for the most important tokens, while maintaining approximate representations for the remaining tokens.
Our method first performs clustering to group together semantically similar key vectors, and then uses the cluster centroids both to identify important key vectors and to approximate the remaining key vectors in order to retain high accuracy.
Additionally, in order to accelerate long generation tasks, we design a fast cluster update process to quickly re-cluster the input and previously generated tokens, thereby allowing for accelerating attention to the previous output tokens.
We evaluate our method using emerging LRMs such as Qwen-8B and Deepseek-R1-Distil-Qwen2.5-14B, demonstrating that our approach can maintain accuracy on complex reasoning tasks even with aggressive attention sparsity settings.
We also provide kernel implementations to demonstrate the practical efficiency gains from our method, achieving up to 4.5$\times$ speedup for attention in long-context reasoning applications. Coleman Hooper, Sebastian Zhao, Luca Manolache, Sehoon Kim 0001, Michael W. Mahoney, Sophia Shao, Kurt Keutzer, Amir Gholami |
NeurIPS | 1 |
| 2024 | SqueezeLLM: Dense-and-Sparse QuantizationabstractGenerative Large Language Models (LLMs) have demonstrated remarkable results for a wide range of tasks. However, deploying these models for inference has been a significant challenge due to their unprecedented resource requirements. This has forced existing deployment frameworks to use multi-GPU inference pipelines, which are often complex and costly, or to use smaller and less performant models. In this work, we demonstrate that the main bottleneck for generative inference with LLMs is memory bandwidth, rather than compute, specifically for single batch inference. While quantization has emerged as a promising solution by representing weights with reduced precision, previous efforts have often resulted in notable performance degradation. To address this, we introduce SqueezeLLM, a post-training quantization framework that not only enables lossless compression to ultra-low precisions of up to 3-bit, but also achieves higher quantization performance under the same memory constraint. Our framework incorporates two novel ideas: (i) sensitivity-based non-uniform quantization, which searches for the optimal bit precision assignment based on second-order information; and (ii) the Dense-and-Sparse decomposition that stores outliers and sensitive weight values in an efficient sparse format. When applied to the LLaMA models, our 3-bit quantization significantly reduces the perplexity gap from the FP16 baseline by up to 2.1x as compared to the state-of-the-art methods with the same memory requirement. Furthermore, when deployed on an A6000 GPU, our quantized models achieve up to 2.3x speedup compared to the baseline. Our code is available at https://github.com/SqueezeAILab/SqueezeLLM. Sehoon Kim 0001, Coleman Hooper, Amir Gholami, Zhen Dong 0003, Xiuyu Li, Sheng Shen 0001, Michael W. Mahoney, Kurt Keutzer |
ICML | 2 |
| 2024 | KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationabstractLLMs are seeing growing use for applications which require large context windows, and with these large context windows KV cache activations surface as the dominant contributor to memory consumption during inference. Quantization is a promising approach for compressing KV cache activations; however, existing solutions fail to represent activations accurately in sub-4-bit precision. Our work, KVQuant, facilitates low precision KV cache quantization by incorporating several novel methods: (i) Per-Channel Key Quantization, where we adjust the dimension along which we quantize the Key activations to better match the distribution; (ii) Pre-RoPE Key Quantization, where we quantize Key activations before the rotary positional embedding to mitigate its impact on quantization; (iii) Non-Uniform KV Cache Quantization, where we derive per-layer sensitivity-weighted non-uniform datatypes that better represent the distributions; and (iv) Per-Vector Dense-and-Sparse Quantization, where we isolate outliers separately for each vector to minimize skews in quantization ranges. By applying our method to the LLaMA, Llama-2, Llama-3, and Mistral models, we achieve < 0.1 perplexity degradation with 3-bit quantization on both Wikitext-2 and C4, outperforming existing approaches. Our method enables serving LLaMA-7B with a context length of up to 1 million on a single A100-80GB GPU and up to 10 million on an 8-GPU system. We develop custom CUDA kernels for KVQuant, showing that we can achieve up to ~1.7x speedups, compared to baseline fp16 matrix-vector multiplications, for the LLaMA-7B model. Coleman Hooper, Sehoon Kim 0001, Hiva Mohammadzadeh, Michael W. Mahoney, Sophia Shao, Kurt Keutzer, Amir Gholami |
NeurIPS | 1 |
| 2021 | SM6: A 16nm System-on-Chip for Accurate and Noise-Robust Attention-Based NLP Applications : The 33rd Hot Chips Symposium - August 22-24, 2021abstractIn this work, we present SM6, an SoC architecture for real-time denoised speech and NLP pipelines, featuring (1) MSSE: an unsupervised probabilistic sound source separation accelerator, (2) FlexNLP: a programmable inference accelerator for attention-based seq2seq DNNs using adaptive floating-point datatypes for wide dynamic range computations, (3) a dual-core Arm Cortex A53 CPU cluster, which provides on-demand SIMD FFT processing, and operating system support. In adverse acoustic conditions, MSSE allows FlexNLP to store up to 6x smaller ASR models obviating the very inefficient strategy of scaling up the DNN model to achieve noise robustness. MSSE and FlexNLP produce efficiency ranges of 4.33-17.6 Gsamples/s/W and 2.6-7.8TFLOPs/W, respectively, with per-frame end-to-end latencies of 15-45ms. Thierry Tambe, En-Yu Yang, Glenn G. Ko, Yuji Chai, Coleman Hooper, Marco Donato, Paul N. Whatmough, Alexander M. Rush, David Brooks 0001, Gu-Yeon Wei |
HCS | 5 |
| 2021 | EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP InferenceabstractTransformer-based language models such as BERT provide significant accuracy improvement to a multitude of natural language processing (NLP) tasks. However, their hefty computational and memory demands make them challenging to deploy to resource-constrained edge platforms with strict latency requirements. Thierry Tambe, Coleman Hooper, Lillian Pentecost, En-Yu Yang, Marco Donato, Victor Sanh, Paul N. Whatmough, Alexander M. Rush, David Brooks 0001, Gu-Yeon Wei |
MICRO | 2 |