EDBT 2026 Demo / reviewers in the wild / expert
Endian Li
dblp:387/1921
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0004-4130-0912ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 61% Memory systems · 30% GPUs and heterogeneous computing · 9% | |
| Artificial intelligence
1 paper |
Language models and text generation · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › computational storage
computational storage device |
0.9 | 1 | 2025 | InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference · HPCA 2025 |
Storage systems
flash and SSD |
0.9 | 1 | 2025 | InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference · HPCA 2025 |
Memory systems › cache management
KV cache management |
0.9 | 1 | 2025 | InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference · HPCA 2025 |
Natural language and speech › Language models and text generation › large language model inference
long-context inference |
0.3 | 1 | 2025 | InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM Inference · HPCA 2025 |
Methods — techniques the papers use, named apart from their topics
p2p transmission · 1.7attention offloading · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SPDK+: Low Latency or High Power Efficiency? We Take BothabstractSPDK, as one of the most efficient I/O storage software, is capable of delivering the lowest I/O latency. Unfortunately, the polling mechanism in SPDK wastes tremendous CPU clock cycles, especially under small I/O operations and low queue depths. Although SPDK supports the conventional interrupt method, it does not improve power efficiency under such circumstances. To address this issue, we propose SPDK+, which enables the user interrupt feature in the SPDK to achieve both low latency and high power efficiency. Specifically, SPDK+ employs user interrupt handling to directly process MSI-X interrupts from SSD devices and utilizes user wait instructions during IO wait periods to conserve power. The comprehensive evaluation results show that SPDK+ achieves up to 49.5% power efficiency improvement while keeping the I/O latency almost unchanged compared with SPDK. Endian Li, Shushu Yi, Qiao Li 0001, Diyu Zhou, Zhenlin Wang 0003, Xiaolin Wang 0001, Bo Mao 0003, Yingwei Luo, Ke Zhou 0001, Jie Zhang 0048 |
HotStorage | 1 |
| 2025 | InstAttention: In-Storage Attention Offloading for Cost-Effective Long-Context LLM InferenceabstractThe widespread of Large Language Models (LLMs) marks a significant milestone in generative AI. Nevertheless, the increasing context length and batch size in offline LLM inference escalate the memory requirement of the key-value (KV) cache, which imposes a huge burden on the GPU VRAM, especially for resource-constrained scenarios (e.g., edge computing). Several cost-effective solutions leverage host memory or SSDs to reduce storage costs for offline inference scenarios and improve the throughput. Nevertheless, they suffer from significant performance penalties imposed by intensive KV cache accesses due to limited PCIe bandwidth. To address these issues, we propose InstAttention, a novel LLM inference system that offloads the most performance-critical computation (i.e., attention in decoding phase) and data (i.e., KV cache) parts to Computational Storage Drives (CSDs), which minimize the enormous KV transfer overheads. InstAttention designs a dedicated flashaware in-storage attention engine with KV cache management mechanisms to exploit the high internal bandwidths of CSDs instead of being limited by the PCIe bandwidth. The optimized P2P transmission between GPU and CSDs further reduces data migration overheads. Experimental results demonstrate that for a 13B model using an NVIDIA A6000 GPU, InstAttention improves throughput for long-sequence inference by up to $11.1 \times$, compared to existing SSD-based solutions such as FlexGen. Xiurui Pan, Endian Li, Qiao Li 0001, Shengwen Liang, Yizhou Shan, Ke Zhou 0001, Yingwei Luo, Xiaolin Wang 0001, Jie Zhang 0048 |
HPCA | 2 |