Endian Li

dblp:387/1921 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0004-4130-0912ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 61% Memory systems · 30% GPUs and heterogeneous computing · 9%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › computational storage
computational storage device
0.912025
InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference · HPCA 2025
Storage systems
flash and SSD
0.912025
InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference · HPCA 2025
Memory systems › cache management
KV cache management
0.912025
InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference · HPCA 2025
Natural language and speech › Language models and text generation › large language model inference
long-context inference
0.312025
InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference · HPCA 2025

Methods — techniques the papers use, named apart from their topics

p2p transmission · 1.7attention offloading · 1.7
YearPublicationVenuePosition
2025 SPDK+: Low Latency or High Power Efficiency? We Take Both
abstract
SPDK, as one of the most efficient I/O storage software, is capable of delivering the lowest I/O latency. Unfortunately, the polling mechanism in SPDK wastes tremendous CPU clock cycles, especially under small I/O operations and low queue depths. Although SPDK supports the conventional interrupt method, it does not improve power efficiency under such circumstances. To address this issue, we propose SPDK+, which enables the user interrupt feature in the SPDK to achieve both low latency and high power efficiency. Specifically, SPDK+ employs user interrupt handling to directly process MSI-X interrupts from SSD devices and utilizes user wait instructions during IO wait periods to conserve power. The comprehensive evaluation results show that SPDK+ achieves up to 49.5% power efficiency improvement while keeping the I/O latency almost unchanged compared with SPDK.
Endian Li, Shushu Yi, Qiao Li 0001, Diyu Zhou, Zhenlin Wang 0003, Xiaolin Wang 0001, Bo Mao 0003, Yingwei Luo, Ke Zhou 0001, Jie Zhang 0048
HotStorage1
2025 InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference
abstract
The widespread of Large Language Models (LLMs) marks a significant milestone in generative AI. Nevertheless, the increasing context length and batch size in offline LLM inference escalate the memory requirement of the key-value (KV) cache, which imposes a huge burden on the GPU VRAM, especially for resource-constrained scenarios (e.g., edge computing). Several cost-effective solutions leverage host memory or SSDs to reduce storage costs for offline inference scenarios and improve the throughput. Nevertheless, they suffer from significant performance penalties imposed by intensive KV cache accesses due to limited PCIe bandwidth. To address these issues, we propose InstAttention, a novel LLM inference system that offloads the most performance-critical computation (i.e., attention in decoding phase) and data (i.e., KV cache) parts to Computational Storage Drives (CSDs), which minimize the enormous KV transfer overheads. InstAttention designs a dedicated flashaware in-storage attention engine with KV cache management mechanisms to exploit the high internal bandwidths of CSDs instead of being limited by the PCIe bandwidth. The optimized P2P transmission between GPU and CSDs further reduces data migration overheads. Experimental results demonstrate that for a 13B model using an NVIDIA A6000 GPU, InstAttention improves throughput for long-sequence inference by up to $11.1 \times$, compared to existing SSD-based solutions such as FlexGen.
Xiurui Pan, Endian Li, Qiao Li 0001, Shengwen Liang, Yizhou Shan, Ke Zhou 0001, Yingwei Luo, Xiaolin Wang 0001, Jie Zhang 0048
HPCA2