Coleman Hooper

dblp:279/6669 · also Coleman Richard Charles Hooper · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-5890-610XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 57% Language models and text generation · 32% Deep learning architectures and training · 12%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 74% Memory systems · 26%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
2.342025
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization · NeurIPS 2024
SqueezeLLM: Dense-and-Sparse Quantization · ICML 2024
EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference · MICRO 2021
Natural language and speech › Language models and text generation › large language model inference
long-context inference
1.932025
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache · ICML 2025
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization · NeurIPS 2024
Squeezed Attention: Accelerating Long Context Length LLM Inference · ACL (1) 2025
Machine learning › Efficient and distributed learning
inference efficiency
1.722025
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache · ICML 2025
Squeezed Attention: Accelerating Long Context Length LLM Inference · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model inference
1.022024
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization · NeurIPS 2024
SqueezeLLM: Dense-and-Sparse Quantization · ICML 2024
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention
0.912025
Multipole Attention for Efficient Long Context Reasoning · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model
large reasoning model
0.912025
Multipole Attention for Efficient Long Context Reasoning · NeurIPS 2025
Natural language and speech › Language models and text generation › chain-of-thought reasoning
long chain-of-thought reasoning
0.912025
Multipole Attention for Efficient Long Context Reasoning · NeurIPS 2025
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention
0.912025
Multipole Attention for Efficient Long Context Reasoning · NeurIPS 2025
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization
0.812024
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression › quantization
non-uniform quantization
0.812024
SqueezeLLM: Dense-and-Sparse Quantization · ICML 2024
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization
0.812024
SqueezeLLM: Dense-and-Sparse Quantization · ICML 2024
Machine learning › Efficient and distributed learning › model compression
quantization
0.812024
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression › sparsity
weight sparsity
0.812024
SqueezeLLM: Dense-and-Sparse Quantization · ICML 2024
Machine learning › Efficient and distributed learning › inference efficiency
energy-efficient inference
0.512021
EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference · MICRO 2021
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
transformer inference
0.512021
EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference · MICRO 2021
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression
0.312025
Multipole Attention for Efficient Long Context Reasoning · NeurIPS 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.212024
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization · NeurIPS 2024
Natural language and speech › Language models and text generation › pre-trained language model
BERT
0.112021
EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference · MICRO 2021

Methods — techniques the papers use, named apart from their topics

weight quantization · 1.7speculative decoding · 1.7KV cache quantization · 1.7kernel implementation · 0.9clustering · 0.9cluster centroid approximation · 0.9attention squeezing · 0.9sensitivity-based quantization · 0.8second-order information · 0.8per-channel quantization · 0.8non-uniform quantization · 0.8dense-and-sparse quantization · 0.8dense-and-sparse decomposition · 0.8custom CUDA kernels · 0.8sentence-level energy optimization · 0.5
YearPublicationVenuePosition
2025 Squeezed Attention: Accelerating Long Context Length LLM Inference
abstract
Coleman Richard Charles Hooper, Sehoon Kim, Hiva Mohammadzadeh, Monishwaran Maheswaran, Sebastian Zhao, June Paik, Michael W. Mahoney, Kurt Keutzer, Amir Gholami. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Coleman Hooper, Sehoon Kim 0001, Hiva Mohammadzadeh, Monishwaran Maheswaran, Sebastian Zhao, June Paik, Michael W. Mahoney, Kurt Keutzer, Amir Gholami
ACL (1)1
2025 QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
abstract
Large Language Models (LLMs) are increasingly being deployed on edge devices for long-context settings, creating a growing need for fast and efficient long-context inference. In these scenarios, the Key-Value (KV) cache is the primary bottleneck in terms of both GPU memory and latency, as the full KV cache must be loaded for each decoding step. While speculative decoding is a widely accepted technique to accelerate autoregressive decoding, existing methods often struggle to achieve significant speedups due to inefficient KV cache optimization strategies and result in low acceptance rates. To address these challenges, we propose a novel self-speculative decoding framework, QuantSpec, where the draft model shares the architecture of the target model but employs a hierarchical 4-bit quantized KV cache and 4-bit quantized weights for acceleration. QuantSpec maintains high acceptance rates ($>$90\%) and reliably provides consistent end-to-end speedups upto $\sim2.5\times$, outperforming other self-speculative decoding methods that use sparse KV cache for long-context LLM inference. QuantSpec also reduces the memory requirements by $\sim 1.3\times$ compared to these alternatives.
Rishabh Tiwari, Haocheng Xi, Aditya Tomar, Coleman Hooper, Sehoon Kim 0001, Maxwell Horton, Mahyar Najibi, Michael W. Mahoney, Kurt Keutzer, Amir Gholami
ICML4
2025 Multipole Attention for Efficient Long Context Reasoning
abstract
Large Reasoning Models (LRMs) have shown promising accuracy improvements on complex problem-solving tasks. While these models have attained high accuracy by leveraging additional computation at test time, they need to generate long chain-of-thought reasoning in order to think before answering, which requires generating thousands of tokens. While sparse attention methods can help reduce the KV cache pressure induced by this long autoregressive reasoning, these methods can introduce errors which disrupt the reasoning process. Our work addresses these challenges by introducing Multipole Attention, which accelerates autoregressive reasoning by only computing exact attention for the most important tokens, while maintaining approximate representations for the remaining tokens. Our method first performs clustering to group together semantically similar key vectors, and then uses the cluster centroids both to identify important key vectors and to approximate the remaining key vectors in order to retain high accuracy. Additionally, in order to accelerate long generation tasks, we design a fast cluster update process to quickly re-cluster the input and previously generated tokens, thereby allowing for accelerating attention to the previous output tokens. We evaluate our method using emerging LRMs such as Qwen-8B and Deepseek-R1-Distil-Qwen2.5-14B, demonstrating that our approach can maintain accuracy on complex reasoning tasks even with aggressive attention sparsity settings. We also provide kernel implementations to demonstrate the practical efficiency gains from our method, achieving up to 4.5$\times$ speedup for attention in long-context reasoning applications.
Coleman Hooper, Sebastian Zhao, Luca Manolache, Sehoon Kim 0001, Michael W. Mahoney, Sophia Shao, Kurt Keutzer, Amir Gholami
NeurIPS1
2024 SqueezeLLM: Dense-and-Sparse Quantization
abstract
Generative Large Language Models (LLMs) have demonstrated remarkable results for a wide range of tasks. However, deploying these models for inference has been a significant challenge due to their unprecedented resource requirements. This has forced existing deployment frameworks to use multi-GPU inference pipelines, which are often complex and costly, or to use smaller and less performant models. In this work, we demonstrate that the main bottleneck for generative inference with LLMs is memory bandwidth, rather than compute, specifically for single batch inference. While quantization has emerged as a promising solution by representing weights with reduced precision, previous efforts have often resulted in notable performance degradation. To address this, we introduce SqueezeLLM, a post-training quantization framework that not only enables lossless compression to ultra-low precisions of up to 3-bit, but also achieves higher quantization performance under the same memory constraint. Our framework incorporates two novel ideas: (i) sensitivity-based non-uniform quantization, which searches for the optimal bit precision assignment based on second-order information; and (ii) the Dense-and-Sparse decomposition that stores outliers and sensitive weight values in an efficient sparse format. When applied to the LLaMA models, our 3-bit quantization significantly reduces the perplexity gap from the FP16 baseline by up to 2.1x as compared to the state-of-the-art methods with the same memory requirement. Furthermore, when deployed on an A6000 GPU, our quantized models achieve up to 2.3x speedup compared to the baseline. Our code is available at https://github.com/SqueezeAILab/SqueezeLLM.
Sehoon Kim 0001, Coleman Hooper, Amir Gholami, Zhen Dong 0003, Xiuyu Li, Sheng Shen 0001, Michael W. Mahoney, Kurt Keutzer
ICML2
2024 KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
abstract
LLMs are seeing growing use for applications which require large context windows, and with these large context windows KV cache activations surface as the dominant contributor to memory consumption during inference. Quantization is a promising approach for compressing KV cache activations; however, existing solutions fail to represent activations accurately in sub-4-bit precision. Our work, KVQuant, facilitates low precision KV cache quantization by incorporating several novel methods: (i) Per-Channel Key Quantization, where we adjust the dimension along which we quantize the Key activations to better match the distribution; (ii) Pre-RoPE Key Quantization, where we quantize Key activations before the rotary positional embedding to mitigate its impact on quantization; (iii) Non-Uniform KV Cache Quantization, where we derive per-layer sensitivity-weighted non-uniform datatypes that better represent the distributions; and (iv) Per-Vector Dense-and-Sparse Quantization, where we isolate outliers separately for each vector to minimize skews in quantization ranges. By applying our method to the LLaMA, Llama-2, Llama-3, and Mistral models, we achieve < 0.1 perplexity degradation with 3-bit quantization on both Wikitext-2 and C4, outperforming existing approaches. Our method enables serving LLaMA-7B with a context length of up to 1 million on a single A100-80GB GPU and up to 10 million on an 8-GPU system. We develop custom CUDA kernels for KVQuant, showing that we can achieve up to ~1.7x speedups, compared to baseline fp16 matrix-vector multiplications, for the LLaMA-7B model.
Coleman Hooper, Sehoon Kim 0001, Hiva Mohammadzadeh, Michael W. Mahoney, Sophia Shao, Kurt Keutzer, Amir Gholami
NeurIPS1
2021 SM6: A 16nm System-on-Chip for Accurate and Noise-Robust Attention-Based NLP Applications : The 33rd Hot Chips Symposium - August 22-24, 2021
abstract
In this work, we present SM6, an SoC architecture for real-time denoised speech and NLP pipelines, featuring (1) MSSE: an unsupervised probabilistic sound source separation accelerator, (2) FlexNLP: a programmable inference accelerator for attention-based seq2seq DNNs using adaptive floating-point datatypes for wide dynamic range computations, (3) a dual-core Arm Cortex A53 CPU cluster, which provides on-demand SIMD FFT processing, and operating system support. In adverse acoustic conditions, MSSE allows FlexNLP to store up to 6x smaller ASR models obviating the very inefficient strategy of scaling up the DNN model to achieve noise robustness. MSSE and FlexNLP produce efficiency ranges of 4.33-17.6 Gsamples/s/W and 2.6-7.8TFLOPs/W, respectively, with per-frame end-to-end latencies of 15-45ms.
Thierry Tambe, En-Yu Yang, Glenn G. Ko, Yuji Chai, Coleman Hooper, Marco Donato, Paul N. Whatmough, Alexander M. Rush, David Brooks 0001, Gu-Yeon Wei
HCS5
2021 EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference
abstract
Transformer-based language models such as BERT provide significant accuracy improvement to a multitude of natural language processing (NLP) tasks. However, their hefty computational and memory demands make them challenging to deploy to resource-constrained edge platforms with strict latency requirements.
Thierry Tambe, Coleman Hooper, Lillian Pentecost, En-Yu Yang, Marco Donato, Victor Sanh, Paul N. Whatmough, Alexander M. Rush, David Brooks 0001, Gu-Yeon Wei
MICRO2