Lingfeng Xiang

dblp:317/2549 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0009-7303-3595ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Scaling Attention Beyond GPUs for LLM Inference
abstract
Scaling inference for large language models is increasingly constrained by limited GPU memory, primarily due to the expanding intermediate states (KV caches) required for long-context generation and multi-user workloads. Once the KV cache exceeds the capacity of high-bandwidth memory, it must be offloaded to host memory and reloaded on demand, a workflow severely bottlenecked by the CPU–GPU interconnect, typically PCIe. Existing approaches exploiting offload KV caches to CPU memory and selectively reload partial segments for attention computation often underutilize CPU compute resources and suffer from accuracy degradation. We present Beyond, a drop-in runtime that integrates a smart offloading scheme to selectively identify and retain salient KV entries across continuous decoding sessions, together with a hybrid CPU–GPU attention mechanism for scalable inference. Beyond executes dense attention over recent KV entries stored in GPU memory while performing parallel, per-head sparse attention on salient contextual KV entries residing in CPU memory. The outputs are fused efficiently through a log-sum-exp scheme. During the bandwidth-constrained decoding phase, oversized KV caches are processed cooperatively by the aggregated CPU and GPU memory bandwidth, with only minimal PCIe data movement. Experiments across diverse models and workloads demonstrate that Beyond improves scalability, supports longer sequences and larger batch sizes, and outperforms existing sparse attention baselines in both efficiency and accuracy—all on commodity GPU hardware.
Weishu Deng, Peiran Du, Lingfeng Xiang, Chen Zhong 0002, Faraz Ahmed, Lianjie Cao, Puneet Sharma 0001, Song Jiang 0001, Hui Lu 0001, Jia Rao
HPDC4
2025 Can Hardware Outsmart Software in Tiered Memory Management? A CMM-H Case Study
abstract
With the advent of Compute Express Link (CXL), hardware-managed memory tiering has become a reality. In this paper, we investigate Samsung's CXL Memory Module-Hybrid (CMM-H), a CXL Type 3 device integrating DRAM and NAND flash managed by an FPGA-based controller and providing byte-addressable memory interface via the cxl.mem protocol. We perform a detailed evaluation of CMM-H and compare its performance with OS-level and block-level tiering solutions. Our results highlight the performance benefits of CMM-H for cache-hit scenarios and identify key limitations for cache-miss situations, offering insights into the trade-offs involved in adopting hardware-managed memory tiering in emerging CXL-based systems.
Lingfeng Xiang, Lianjie Cao, Faraz Ahmed, Jia Rao, Hui Lu 0001, Puneet Sharma 0001
SYSTOR3
2025 DSA-2LM: A CPU-Free Tiered Memory Architecture with Intel DSA
Ruili Liu, Teng Ma 0006, Yingdi Shan, Zheng Liu 0022, Lingfeng Xiang, Hui Lu 0001, Jia Rao, Kang Chen 0001, Yongwei Wu 0001
USENIX ATC7
2024 Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration
Lingfeng Xiang, Weishu Deng, Hui Lu 0001, Jia Rao, Ren Wang 0001
OSDI1
2023 P2CACHE: Exploring Tiered Memory for In-Kernel File Systems Caching
Lingfeng Xiang, Jia Rao, Hui Lu 0001
USENIX ATC2
2022 Characterizing the performance of intel optane persistent memory: a close look at its on-DIMM buffering
abstract
We present a comprehensive and in-depth study of Intel Optane DC persistent memory (DCPMM). Our focus is on exploring the internal design of Optane's on-DIMM read-write buffering and its impacts on application-perceived performance, read and write amplifications, the overhead of different types of persists, and the tradeoffs between persistency models. While our measurements confirm the results of the existing profiling studies, we have new discoveries and offer new insights. Notably, we find that read and write are managed differently in separate on-DIMM read and write buffers. Comparable in size, the two buffers serve distinct purposes. The read buffer offers higher concurrency and effective on-DIMM prefetching, leading to high read bandwidth and superior sequential performance. However, it does not help hide media access latency. In contrast, the write buffer offers limited concurrency but is a critical stage in a pipeline that supports asynchronous write in the DDR-T protocol. Surprisingly, in addition to write coalescing, the write buffer delivers lower than read and consistent write latency regardless of the working set size, the type of write, the access pattern, or the persistency model. Furthermore, we discover that the mismatch between cacheline access granularity and the 3D-Xpoint media access granularity negatively impacts the effectiveness of CPU cache prefetching and leads to wasted persistent memory bandwidth.
Lingfeng Xiang, Xingsheng Zhao, Jia Rao, Song Jiang 0001, Hong Jiang 0001
EuroSys1