VLDB 2026 Research / reviewers in the wild / expert
Erge Xiang
dblp:375/7258
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0005-2963-2479ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 46% Memory systems · 46% GPUs and heterogeneous computing · 7% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › model compression
lookup-table-based inference |
1.0 | 1 | 2026 | MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware · IEEE Trans. Computers 2026 |
Machine learning › Efficient and distributed learning
model compression |
1.0 | 1 | 2026 | MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware · IEEE Trans. Computers 2026 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator |
1.0 | 1 | 2026 | MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware · IEEE Trans. Computers 2026 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.0 | 1 | 2026 | MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware · IEEE Trans. Computers 2026 |
Memory systems › processing-in-memory
near-bank processing-in-memory |
1.0 | 1 | 2026 | MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware · IEEE Trans. Computers 2026 |
Memory systems
processing-in-memory |
1.0 | 1 | 2026 | MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware · IEEE Trans. Computers 2026 |
Methods — techniques the papers use, named apart from their topics
quantization · 2.0model partitioning · 2.0lookup table · 2.0learning-based LUT construction · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory HardwareabstractLarge language models (LLMs) are prominent for their superior ability in language understanding and generation. However, a notorious problem for LLM inference is low computational utilization caused by the memory bottleneck, since it typically requires large memory capacity and high bandwidth to process neural weights. By integrating processing cores into memory, Processing-In-Memory (PIM) architecture excels at alleviating memory bottleneck; with the recent release of the first commodity near-bank PIM hardware (NBP), PIM becomes off-the-shelf and shows great potential for accelerating LLM inference practically. However, simply shoehorning LLM inference on NBP can not achieve satisfactory performance due to its inherent limitations: weak compute performance, frequent cache misses caused by the limited working memory capacity, and poor inter-PIM-core communication bandwidth. To address these limitations, we propose MI-LLM, an efficient system deploying LLM inference on NBP hardware. Its key idea is to build NBP-aware Lookup Tables (LUTs) and completely replace multiplications with lookups on LUTs, thereby mitigating the limitation of weak compute performance. 1) To reduce the model accuracy drop caused by the use of LUT, MI-LLM tailors a learning-based LUT construction method to maintain the model accuracy. 2) To cope with frequent cache misses caused by LUT sizes far exceeding PIM working memory capacity, MI-LLM introduces the design of PIM-aware linear kernel, with the optimization of intra-row and inter-row reordering enabled, to enhance LUT lookup locality. 3) MI-LLM further proposes a model partitioning scheme to minimize inter-PIM-core communication. Kernel-level benchmarks reveal that MI-LLM achieves a 9% throughput improvement and an 11% increase in energy efficiency over GPU implementations. Compared to FP8 quantization, MI-LLM incurs only a 0.24 times increase in perplexity, demonstrating minimal accuracy degradation. Moreover, in our end-to-end evaluation, MI-LLM requires 80% fewer ALU operation ticks per output token than the GPU baseline. Puyun Hu, Minhui Xie, Linjiang Li, Kuiyaohui Zhang, Erge Xiang, Jing Wang 0055, Size Zheng 0001, Xiao Zhang 0001, Yunpeng Chai |
IEEE Trans. Computers | 5 |