VLDB 2026 Research / reviewers in the wild / expert
Yayue Hou
dblp:358/2500
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0009-6780-076XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STARC: Selective Token Access with Remapping and Clustering for Efficient LLM Decoding on PIM SystemsabstractServing large language models (LLMs) places significant pressure on memory systems due to frequent accesses and growing key–value (KV) caches as context lengths increase. Processing-in-memory (PIM) architectures offer high internal bandwidth and near-data compute parallelism, but current designs target dense attention and perform poorly under the irregular access patterns of dynamic KV cache sparsity. To mitigate this limitation, we propose STARC, a sparsity-optimized data mapping scheme for efficient LLM decoding on PIM. STARC clusters semantically similar KV pairs and co-locates them contiguously within PIM banks, enabling retrieval at cluster granularity by matching queries against precomputed centroids. This bridges the gap between fine-grained sparse attention and row-level PIM operations, improving utilization while minimizing overhead. On a simulated HBM-PIM system, under constrained KV budgets, STARC achieves up to 78% and 65% reductions in attention-layer latency and energy over token-wise sparsity methods, and up to 93% and 92% reductions relative to full attention, while preserving model accuracy. Zehao Fan, Yunzhen Liu, Garrett Gagnon, Yayue Hou, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu 0017 |
ASPLOS (2) | 5 |
| 2025 | NORA: Noise-Optimized Rescaling of LLMs on Analog Compute-in-Memory AcceleratorsabstractLarge Language Models (LLMs) have become critical in AI applications, yet current digital AI accelerators suffer from significant energy inefficiencies due to frequent data movement. Analog compute-in-memory (CIM) accelerators offer a potential solution for improving energy efficiency but introduce non-idealities that can degrade LLM accuracy. While analog CIM has been extensively studied for traditional deep neural networks, its impact on LLMs remains unexplored, particularly concerning the large influence of Analog CIM non-idealities. In this paper, we conduct a sensitivity analysis on the effects of analog-induced noise on LLM accuracy. We find that while LLMs demonstrate robustness to weight-related noise, they are highly sensitive to quantization noise and additive Gaussian noise. Based on these insights, we propose a noise-optimized rescaling method to mitigate LLM accuracy loss by shifting the non-ideality burden from the sensitive input/output to the more resilient weight. Through rescaling, we can implement the OPT-6.7b model on simulated analog CIM hardware with less than 1% accuracy loss from the floating-point baseline, compared to a much higher loss of around 30% without rescaling. Yayue Hou, Hsinyu Tsai, Kaoutar El Maghraoui, Tayfun Gokmen, Geoffrey W. Burr, Liu Liu 0017 |
DATE | 1 |
| 2025 | SAGE: Saliency-Aware Grouping for Efficient Mapping of LLMs on Analog Compute-in-MemoryabstractLarge Language Models (LLMs) demand high memory bandwidth and computational efficiency, posing significant challenges for deployment on traditional digital accelerators. Analog Compute-in-Memory (ACIM) architectures offer an attractive alternative by co-locating storage and computation to reduce data movement. However, executing LLMs on ACIM systems remains challenging due to hardware non-idealities and the unique statistical properties of LLM inputs and outputs in FC layers. In particular, long-tailed data distributions containing large-amplitude "salient values" degrade analog signal quality under quantization and system noise. In this work, we propose SAGE (Saliency-Aware Grouping for Efficient Mapping), a training-free strategy that improves noise resilience by reordering weight and input channels of FC layers based on statistical characteristics of LLMs. We identify kurtosis as a key factor affecting analog robustness and develop a saliency-aware mapping method that reduces output kurtosis to enhance the signal-to-noise ratio. We further introduce a reconfigurable tile design that supports mixed-precision execution and maximizes array utilization across layers. Evaluations on multiple LLMs and benchmarks show that SAGE significantly improves inference accuracy and energy efficiency based on ACIM simulation without requiring retraining. Yayue Hou, Garrett Gagnon, Hsinyu Tsai, Kaoutar El Maghraoui, Geoffrey W. Burr, Liu Liu 0017 |
ICCAD | 1 |