Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Erge Xiang

dblp:375/7258 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0005-2963-2479ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 46% Memory systems · 46% GPUs and heterogeneous computing · 7%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › model compression
lookup-table-based inference
1.012026
MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware · IEEE Trans. Computers 2026
Machine learning › Efficient and distributed learning
model compression
1.012026
MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware · IEEE Trans. Computers 2026
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator
1.012026
MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware · IEEE Trans. Computers 2026
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.012026
MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware · IEEE Trans. Computers 2026
Memory systems › processing-in-memory
near-bank processing-in-memory
1.012026
MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware · IEEE Trans. Computers 2026
Memory systems
processing-in-memory
1.012026
MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware · IEEE Trans. Computers 2026

Methods — techniques the papers use, named apart from their topics

quantization · 2.0model partitioning · 2.0lookup table · 2.0learning-based LUT construction · 2.0
YearPublicationVenuePosition
2026 MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware
abstract
Large language models (LLMs) are prominent for their superior ability in language understanding and generation. However, a notorious problem for LLM inference is low computational utilization caused by the memory bottleneck, since it typically requires large memory capacity and high bandwidth to process neural weights. By integrating processing cores into memory, Processing-In-Memory (PIM) architecture excels at alleviating memory bottleneck; with the recent release of the first commodity near-bank PIM hardware (NBP), PIM becomes off-the-shelf and shows great potential for accelerating LLM inference practically. However, simply shoehorning LLM inference on NBP can not achieve satisfactory performance due to its inherent limitations: weak compute performance, frequent cache misses caused by the limited working memory capacity, and poor inter-PIM-core communication bandwidth. To address these limitations, we propose MI-LLM, an efficient system deploying LLM inference on NBP hardware. Its key idea is to build NBP-aware Lookup Tables (LUTs) and completely replace multiplications with lookups on LUTs, thereby mitigating the limitation of weak compute performance. 1) To reduce the model accuracy drop caused by the use of LUT, MI-LLM tailors a learning-based LUT construction method to maintain the model accuracy. 2) To cope with frequent cache misses caused by LUT sizes far exceeding PIM working memory capacity, MI-LLM introduces the design of PIM-aware linear kernel, with the optimization of intra-row and inter-row reordering enabled, to enhance LUT lookup locality. 3) MI-LLM further proposes a model partitioning scheme to minimize inter-PIM-core communication. Kernel-level benchmarks reveal that MI-LLM achieves a 9% throughput improvement and an 11% increase in energy efficiency over GPU implementations. Compared to FP8 quantization, MI-LLM incurs only a 0.24 times increase in perplexity, demonstrating minimal accuracy degradation. Moreover, in our end-to-end evaluation, MI-LLM requires 80% fewer ALU operation ticks per output token than the GPU baseline.
Puyun Hu, Minhui Xie, Linjiang Li, Kuiyaohui Zhang, Erge Xiang, Jing Wang 0055, Size Zheng 0001, Xiao Zhang 0001, Yunpeng Chai
IEEE Trans. Computers5