EDBT 2026 Demo / reviewers in the wild / expert
Jonah Yi
dblp:355/5603 · also Jonah Wonkyu Yi
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Processor architecture and microarchitecture · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
inference efficiency |
1.5 | 2 | 2024 | NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention · NeurIPS 2024 KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
attention computation |
0.8 | 1 | 2024 | NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression › quantization
KV cache quantization |
0.8 | 1 | 2024 | KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model compression › quantization
low-bit quantization |
0.8 | 1 | 2024 | KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization · NeurIPS 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization · NeurIPS 2024 |
Processor architecture and microarchitecture
SIMD |
0.8 | 1 | 2024 | NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › model inference
transformer inference |
0.5 | 2 | 2024 | NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free Attention · NeurIPS 2024 KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled Quantization · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
SIMD lookup · 1.54-bit quantization · 1.5entropy analysis · 0.8coupled quantization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | KV Cache is 1 Bit Per Channel: Efficient Large Language Model Inference with Coupled QuantizationabstractEfficient deployment of Large Language Models (LLMs) requires batching multiple requests together to improve throughput. As batch size, context length, or model size increases, the size of key and value (KV) cache quickly becomes the main contributor to GPU memory usage and the bottleneck of inference latency and throughput. Quantization has emerged as an effective technique for KV cache compression, but existing methods still fail at very low bit widths. Currently, KV cache quantization is performed per-channel or per-token independently. Our analysis shows that distinct channels of a key/value activation embedding are highly interdependent, and the joint entropy of multiple channels grows at a slower rate than the sum of their marginal entropy, which implies that per-channel independent quantization is sub-optimal. To mitigate this sub-optimality, we propose Coupled Quantization (CQ), which couples multiple key/value channels together for quantization to exploit their interdependence and encode the activations in a more information-efficient manner. Extensive experiments reveal that CQ compares favorably with existing baselines in preserving model quality, and improves inference throughput by 1.4–3.5$\times$ relative to the uncompressed baseline. Furthermore, we demonstrate that CQ can preserve model quality reasonably with KV cache quantized down to 1 bit. Tianyi Zhang 0011, Jonah Yi, Zhaozhuo Xu, Anshumali Shrivastava |
NeurIPS | 2 |
| 2024 | NoMAD-Attention: Efficient LLM Inference on CPUs Through Multiply-add-free AttentionabstractLarge Language Model (LLM) inference on Central Processing Units (CPU) is challenging due to the vast quantities of Multiply-Add (MAD) matrix operations in the attention computations. This paper highlights a rare gem in modern CPUs, Single-Instruction-Multiple-Data (SIMD) registers, which allows for ultra-low-latency lookups in a batch. We leverage this unique capability to propose NoMAD-Attention, an efficient attention algorithm that replaces MAD operations with in-register lookups. Through hardware-aware algorithmic designs, NoMAD-Attention achieves the computation of attention scores using repeated fast accesses to SIMD registers. NoMAD-Attention works with pre-trained attention-based LLMs without model finetuning. Extensive empirical evaluations demonstrate that NoMAD-Attention maintains the quality of the original LLMs well and speeds up the 4-bit quantized LLaMA-7B-based model by up to $2 \times$ at 16k context length. Tianyi Zhang 0011, Jonah Yi, Zhaozhuo Xu, Anshumali Shrivastava |
NeurIPS | 2 |