Jiangpeng Li

dblp:01/9848 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Five-Minute Rule 40 Years Later: A First-Principles Revisit for Modern Memory Hierarchy
abstract
In 1987, Jim Gray and Gianfranco Putzolu introduced the five-minute rule, a simple, storage-memory-economics-based heuristic for deciding when data should live in DRAM rather than on storage. Subsequent revisits to the rule largely retained that economics-only view, leaving host costs, feasibility limits, and workload behavior out of scope. This paper revisits the rule from first principles, integrating host costs, DRAM bandwidth/capacity, and physics-grounded models of SSD performance and cost, and then embedding these elements in a constraint- and workload-aware framework that yields actionable provisioning guidance. We show that, for modern AI platforms, especially GPU-centric hosts paired with ultra-high-IOPS SSDs engineered for fine-grained random access, the DRAM$\leftrightarrow$flash caching threshold collapses from minutes to a few seconds. This shift reframes NAND flash memory as an \emph{active data tier} and exposes a broad research space across the hardware-software stack. We further introduce MQSim-Next, a calibrated SSD simulator that supports validation and sensitivity analysis and facilitates future architectural and system research. Finally, we present two concrete case studies that showcase the software system design space opened by such memory hierarchy paradigm shift. Overall, we turn a classical heuristic into an actionable, feasibility-aware analysis and provisioning framework and set the stage for further research on AI-era memory hierarchy.
Tong Zhang 0002, Vikram S. Mailthody, Linsen Ma, Chris J. Newburn, Teresa Zhang, Jiangpeng Li, Hao Zhong 0006, Wen-Mei W. Hwu
ISCA8
2025 Parallax-robust correlation volume for optical flow computation neural networks
Jiangpeng Li, Yan Niu
Multim. Syst.1
2024 Prioritized Planning for Large-Scale Multiple-AGV Scheduling Problem in Smart Manufacturing
abstract
Robotics and automation is one of crucial trend in smart manufacturing to improve production efficiency. Au-tomated guided vehicles (AGVs) are a type of mobile robot used for material handling and have become widely utilized to achieve transportation automation. The usage of multiple AGVs introduces potential risks, such as traffic conflicts and safety risk. To meet high production demands, numerous shop floors are set up for large-scale manufacturing. Thus, reasonable and efficient AGV scheduling is vital for real-world operations. This paper proposes an efficient and scalable prioritized planning algorithm for large-scale multiple-AGV scheduling problem in manufacturing. The algorithm sequentially addresses two primary sub-problems: job assignment and conflict-free routing. The results of job assignment dictate the routes taken by the AGVs. In job assignment, jobs are allocated sequentially based on their pickup times. In conflict-free routing, AGV priorities are predefined, ensuring that higher priority AGVs maintain their movement while adjustments are made only to lower priority AGV plans when conflicts arise. Simulation is conducted on two real shop floor layouts and demonstrates the effectiveness and high efficiency of proposed algorithm. Even in a large-scale layout with 500 jobs and 20 AGVs, the computation time is only around 21 seconds.
Jiarong Yao, Jiangpeng Li, Rong Su 0001, Keck Voon Ling
ICARCV3
2023 Elastic RAID: Implementing RAID over SSDs with Built-in Transparent Compression
abstract
This paper studies how RAID (redundant array of independent disks) could take full advantage of modern SSDs (solid-state drives) with built-in transparent compression. In current practice, RAID users are forced to choose a specific RAID level (e.g., RAID 10 or RAID 5) with a fixed storage cost vs. speed performance trade-off. The commercial market is witnessing the emergence of a new family of SSDs that can internally perform hardware-based lossless compression on each 4KB LBA (logical block address) block, transparent to host OS and user applications. Beyond straightforwardly reducing the RAID storage cost, such modern SSDs make it possible to relieve RAID users from being locked into a fixed storage cost vs. speed performance trade-off. In particular, RAID systems could opportunistically leverage higher-than-expected runtime user data compressibility to enable dynamic RAID level conversion to improve the speed performance without compromising the effective storage capacity. This paper presents techniques to enable and optimize the practical implementation of such elastic RAID systems. We implemented a Linux software-based elastic RAID prototype that supports dynamic conversion between RAID 5 and RAID 10. Compared with a baseline software-based RAID 5, under sufficient runtime data compressibility that enables the conversion from RAID 5 to RAID 10 over 60% of user data, the elastic RAID could improve the 4KB random write IOPS (I/O per second) by 42% and 4KB random read IOPS in degraded mode by 46%, while maintaining the same effective storage capacity.
Jiangpeng Li, Yang Liu 0256, Tong Zhang 0002
SYSTOR2
2022 Closing the B+-tree vs. LSM-tree Write Amplification Gap on Modern Storage Hardware with Built-in Transparent Compression
Yifan Qiao 0003, Xubin Chen, Jiangpeng Li, Yang Liu 0256, Tong Zhang 0002
FAST4
2021 KallaxDB: A Table-less Hash-based Key-Value Store on Storage Hardware with Built-in Transparent Compression
abstract
This paper studies the design of a key-value (KV) store that can take full advantage of modern storage hardware with built-in transparent compression capability. Many modern storage appliances/drives implement hardware-based data compression, transparent to OS and applications. Moreover, the growing deployment of hardware-based compression in Cloud infrastructure leads to the imminent arrival of Cloud-based storage hardware with built-in transparent compression. By decoupling the logical storage space utilization efficiency from the true physical storage usage, transparent compression allows data management software to purposely waste logical storage space in return for simpler data structures and algorithms, leading to lower implementation complexity and higher performance. This work proposes a table-less hash-based KV store, where the basic idea is to hash the key space directly onto the logical storage space without using a hash table at all. With a substantially simplified data structure, this approach is subject to significant logical storage space under-utilization, which can be seamlessly mitigated by storage hardware with transparent compression. This paper presents the basic KV store architecture, and develops mathematical formulations to assist its configuration and analysis. We implemented such a KV store KallaxDB and carried out experiments on a commercial SSD with built-in transparent compression. The results show that, while consuming very little memory resource, it compares favorably with the other modern KV stores in terms of throughput, latency, and CPU usage.
Xubin Chen, Shukun Xu, Yifan Qiao 0003, Yang Liu 0256, Jiangpeng Li, Tong Zhang 0002
DaMoN6
2021 Improving Relational Database Upon the Arrival of Storage Hardware with Built-in Transparent Compression
abstract
This paper presents an approach to enable relational database take full advantage of modern storage hardware with built-in transparent compression. Advanced storage appliances (e.g., all-flash array) and some latest SSDs (solid-state drives) can perform hardware-based data compression, transparently from OS and applications. Moreover, the growing deployment of hardware-based compression capability in Cloud storage infrastructure leads to the imminent arrival of cloud-based storage hardware with built-in transparent compression. To make relational database better leverage modern storage hardware, we propose to deploy a dual in-memory vs. on-storage page format: While pages in database cache memory retain the conventional row-based format, each page on storage devices has a column-based format so that it can be better compressed by storage hardware. We present design techniques that can further improve the on-storage page data compressibility through additional light-weight column data transformation. We the impact of compression algorithms on the selection of column data transformation techniques. We integrated the design techniques into MySQL/InnoDB by adding only about 600 lines of code, and ran Sysbench OLTP workloads on a commercial SSD with built-in transparent compression. The results show that the proposed solution can bring up to 45% additional reduction on the storage cost at only a few percentage of performance degradation.
Yifan Qiao 0003, Xubin Chen, Jingpeng Hao, Jiangpeng Li, Qi Wu 0006, Jingqiang Wang, Yang Liu 0256, Tong Zhang 0002
NAS4
2020 Re-think Data Management Software Design Upon the Arrival of Storage Hardware with Built-in Transparent Compression
Xubin Chen, Jiangpeng Li, Qi Wu 0006, Yang Liu 0256, Hao Zhong 0006, Tong Zhang 0002
HotStorage3
2017 Realizing Transparent OS/Apps Compression in Mobile Devices at Zero Latency Overhead
abstract
Motivated by the significant storage footprint of OS/Apps in mobile devices, this paper studies the realization of OS/Apps transparent compression. In spite of its obvious advantage, this feature however is not widely available in commercial mobile devices, which is due to the justifiable concern on the read latency penalty. In conventional practice on implementing transparent compression, read latency overhead comes from two aspects, including read amplification and decompression computational latency. This paper presents simple yet effective design solutions to eliminate the read amplification at the filesystem level and eliminate the computational latency overhead at the computer architecture level. To demonstrate its practical feasibility, we first implemented a prototyping filesystem to empirically verify the realization of transparent compression with zero read amplification. We further demonstrated that the OS/Apps footprint can be reduced by up to 39 percent on a Nexus 7 tablet installed with Android 5.0. Through application-specific integrated circuit (ASIC) synthesis, we show that the proposed computer architecture level design solution can eliminate the decompression latency overhead with very small silicon cost.
Jiangpeng Li, Hao Wang 0042, Danni Xiong, Jerry Qu, Hyunsuk Shin, Jung Pill Kim, Tong Zhang 0002
IEEE Trans. Computers2
2016 Reducing Solid-State Storage Device Write Stress through Opportunistic In-place Delta Compression
Jiangpeng Li, Hao Wang 0042, Kai Zhao 0005, Tong Zhang 0002
FAST2
2015 How Much Can Data Compressibility Help to Improve NAND Flash Memory Lifetime?
Jiangpeng Li, Kai Zhao 0005, Jun Ma 0012, Tong Zhang 0002
FAST1
2015 Leveraging Progressive Programmability of SLC Flash Pages to Realize Zero-overhead Delta Compression for Metadata Storage
Jiangpeng Li, Kai Zhao 0005, Hao Wang 0042, Tong Zhang 0002
HotStorage2
2015 True-Damage-Aware Enumerative Coding for Improving nand Flash Memory Endurance
abstract
This brief presents a technique that can fully exploit the data dependency of flash memory cell damage to improve the program/erase (P/E) cycling endurance of nand flash memory. The key is to opportunistically leverage data lossless compressibility and utilize the compression gain to realize memory-damage-aware data manipulation to reduce the cycling-induced physical damage. Based upon experiments using commercial sub-22-nm MLC nand flash memory chips, we show that the proposed design technique can improve the P/E cycling endurance by 50%. We further carried out application-specific integrated circuit design to demonstrate the practical feasibility for implementing the proposed design technique.
Jiangpeng Li, Kai Zhao 0005, Jun Ma 0012, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Optimizing the Use of STT-RAM in SSDs Through Data-Dependent Error Tolerance
abstract
This brief presents a design strategy for spin-transfer torque (STT)-RAM to reduce the error-tolerance redundancy overhead and increase effective storage capacity without sacrificing its reliability. The key is to cohesively exploit the run-time data characteristics (e.g., access unit length and access frequency) and the fundamental read disturbance versus sensing error tradeoff in STT-RAM. It presents three specific data-dependent error-tolerance design techniques, and demonstrates their effectiveness in the context of using STT-RAM to replace DRAM in solid-state drives. Based on detailed modeling/simulations down to 22-nm node, we showed that these design solutions can increase the effective STT-RAM storage capacity by 26%, compared with conventional design practice.
Hao Wang 0042, Kai Zhao 0005, Jiangpeng Li, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.3
2014 Over-clocked SSD: Safely running beyond flash memory chip I/O clock specs
abstract
This paper presents a design strategy that enables aggressive use of flash memory chip I/O link over-clocking in solid-state drives (SSDs) without sacrificing storage reliability. The gradual wear-out and process variation of NAND flash memory makes the worst-case oriented error correction code (ECC) in SSDs largely under-utilized most of the time. This work proposes to opportunistically leverage under-utilized error correction strength to allow error-prone flash memory I/O link over-clocking. Its rationale and key design issues are presented and studied in this paper, and its potential effectiveness has been verified through hardware experiments and system simulations. Using sub-22nm NAND flash memory chips with I/O specs of 166MBps, we carried out extensive experiments and show that the proposed design strategy can enable SSDs safely operate with error-prone I/O link running at 275MBps. Trace-driven SSD simulations over a variety of workload traces show the system read response time can be reduced by over 20%.
Kai Zhao 0005, Kalyana S. Venkataraman, Jiangpeng Li, Tong Zhang 0002
HPCA4
2014 Proximate control stream assisted video transcoding for heterogeneous content delivery network
abstract
Video transcoding can be used to facilitate video streaming in content delivery network. A concept of control stream assisted transcoding has been recently proposed aiming to reduce transcoding computational complexity at the cost of data storage/transmission overhead, and the control stream is obtained by directly removing the residual information from the target video bitstream of transcoding. However, it is subject to relatively significant storage/transmission overhead, and more importantly does not well match to the increasingly heterogeneous networking environment with varying computation/storage/transmission resources at different nodes. This work presents a proximate control stream assisted transcoding design strategy that can reduce the control stream size and enable a large storage/transmission vs. computational complexity trade-off design space. Experiments demonstrate its effectiveness and noticeable advantages over other alternatives including simulcast and SVC (scalable video coding).
Jiangpeng Li, Kai Zhao 0005, Tong Zhang 0002
ICIP3
2013 A memory efficient parallel layered QC-LDPC decoder for CMMB systems
Jiangpeng Li, Jun Ma 0012, Guanghui He 0002
Integr.1
2011 Memory efficient layered decoder design with early termination for LDPC codes
abstract
Layered structure is widely used in the design of Low-Density Parity-Check (LDPC) code decoders due to its fast convergence speed. However, correct checking process is difficult to implement in layered decoder, which results in unnecessary iterations. In this paper, an early termination strategy is presented for layered LDPC decoder to avoid redundant number of iterations. This approach makes use of the comparison between current log-likelyhood ratios (LLRs) and updated LLRs of all variable nodes to determine termination criteria of iterations. Furthermore, a non-uniform quantization scheme and an extrinsic messages memory optimization scheme are developed for memory savings. Based on these proposed methods, an LDPC decoder for the Chinese digital mobile TV applications is implemented using a SMIC 130nm CMOS process. The decoder consumes only 171 Kbits memory while achieving 267Mbps for code rate 1/2, and 401Mbps for code rate 3/4.
Jiangpeng Li, Guanghui He 0002, Hexi Hou, Zhejun Zhang, Jun Ma 0012
ISCAS1