Linsen Ma

dblp:349/0074 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0000-8535-7911ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Five-Minute Rule 40 Years Later: A First-Principles Revisit for Modern Memory Hierarchy
abstract
In 1987, Jim Gray and Gianfranco Putzolu introduced the five-minute rule, a simple, storage-memory-economics-based heuristic for deciding when data should live in DRAM rather than on storage. Subsequent revisits to the rule largely retained that economics-only view, leaving host costs, feasibility limits, and workload behavior out of scope. This paper revisits the rule from first principles, integrating host costs, DRAM bandwidth/capacity, and physics-grounded models of SSD performance and cost, and then embedding these elements in a constraint- and workload-aware framework that yields actionable provisioning guidance. We show that, for modern AI platforms, especially GPU-centric hosts paired with ultra-high-IOPS SSDs engineered for fine-grained random access, the DRAM$\leftrightarrow$flash caching threshold collapses from minutes to a few seconds. This shift reframes NAND flash memory as an \emph{active data tier} and exposes a broad research space across the hardware-software stack. We further introduce MQSim-Next, a calibrated SSD simulator that supports validation and sensitivity analysis and facilitates future architectural and system research. Finally, we present two concrete case studies that showcase the software system design space opened by such memory hierarchy paradigm shift. Overall, we turn a classical heuristic into an actionable, feasibility-aware analysis and provisioning framework and set the stage for further research on AI-era memory hierarchy.
Tong Zhang 0002, Vikram S. Mailthody, Linsen Ma, Chris J. Newburn, Teresa Zhang, Jiangpeng Li, Hao Zhong 0006, Wen-Mei W. Hwu
ISCA4
2026 Towards Encrypted Data Compression with Computational Storage Drives
abstract
Modern data center systems need to achieve several critical goals in security, performance, and cost efficiency. However, realizing these goals simultaneously is highly challenging. In secure data storage systems, a common practice is to first compress and encrypt data on the host side and then transmit it to the storage system using a log-based structure. This approach, unfortunately, leads to increased complexity and performance penalty. As an emerging storage technology, Computational Storage Drives (CSD) can not only offload heavy computation burdens to storage device hardware, but also provide a virtualized logical storage space, creating new optimization opportunities. In this paper, we showcase two unique opportunities enabled by the new CSD technology in data storage management. By replacing ordinary SSDs with CSDs, we can realize efficient one-to-one mapping from host-side blocks to storage-side CSD blocks, eliminating the need for a complex log-based structure and the associated heavy-cost operations, such as garbage collections (GC). Moreover, with a carefully redesigned data format in each compression unit, CSDs can transparently remove redundant data across encrypted snapshots in data-intensive environments, such as databases. We have developed a prototype and conducted experiments on ScaleFlux’s CSD 3000 devices to demonstrate the efficacy of these solutions. We hope that our system investigations in this work provide valuable insight into CSDs and inspire researchers and practitioners to explore additional cases for adopting CSDs to improve the performance and productivity of data center systems.
Linsen Ma, Rui Xie 0006, Feng Chen 0005, Xiaodong Zhang 0001, Tong Zhang 0002
SSDBM1
2026 TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
abstract
LLM inference is increasingly limited by memory bandwidth, and the bottleneck worsens at long context as the KV cache grows. CXL memory adds capacity to offload weights and KV, but its link and device-side DDR bandwidth are far below HBM, so decoding stalls once traffic shifts to the CXL tier. Many CXL controllers are starting to add genericlosslesscompression, yet applying commodity codecs directly to standard word-major LLM tensors is largely ineffective, especially for token-major KV streams. We propose TRACE (Traffic-Reduced Architecture for Compression and Elasticity), which preserves the unmodified CXL.mem interface but changes the device-internal representation. It stores tensors in a channel-major, disaggregated bit-plane layout, and applies a KV-specific transform before compression, converting mixed-field words into low-entropy plane streams that commodity codecs can compress. The same substrate enables precision-proportional fetch by reading only the required bit-planes. Across public LLMs, TRACE reduces BF16 weight footprint by 25.2% and BF16 KV footprint by 46.9% losslessly, with per-layer KV ratios peaking at 2.69×. In tracedriven system modeling, once KV spills to CXL, GPT-OSS-120B-MXFP4 improves throughput at 128k tokens from 16.28 to 68.99 tok/s (4.24×). DRAMSim3 shows up to 40.3% lower DRAM access energy under plane-aligned fetch. A 7nm SystemVerilog implementation sustains 256 GB/s device bandwidth. Relative to a CXL controller with generic inline lossless compression, TRACE only adds 7.2% area, 4.7% power, and 6.0% load-to-use latency at 2 GHz and 0.7V.
Rui Xie 0006, Asad Ul Haq, Yunhua Fang, Linsen Ma, Zirak Burzin Engineer, Liu Liu 0017, Tong Zhang 0002
IEEE Trans. Computers4
2023 ZipKV: In-Memory Key-Value Store with Built-In Data Compression
abstract
This paper studies how to mitigate the speed performance loss caused by integrating block data compression into in-memory key-value ‍(KV) stores. Despite extensive prior research on in-memory KV stores, little focus has been given to memory usage reduction via block data compression (e.g., LZ4, ZSTD) due to potential performance degradation. This paper introduces design techniques to mitigate compression-induced performance degradation by utilizing decompression streaming, latency differences between compression and decompression, and data access locality in real-world workloads. These techniques can be incorporated into conventional hash or B+-tree indexing structures, enabling integration with most in-memory KV stores without altering their core indexing data structures. For demonstration, we implemented ZipKV that incorporates the developed design techniques. Compared with RocksDB ‍(in-memory mode) that employs the log-structured merge tree indexing data structure with natural support of block data compression, ZipKV realizes similar memory usage reduction via block data compression, reduces the point query latency by 68% ‍(LZ4) and 58% ‍(ZSTD), and achieves up to 3.8× ‍(LZ4) and 2.7× ‍(ZSTD) point query throughput.
Linsen Ma, Rui Xie 0006, Tong Zhang 0002
ISMM1