EDBT 2026 Demo / reviewers in the wild / expert
Renping Liu 0002
dblp:224/5679-2
· DBLP profile ↗
20ranked-venue papers
7as first author
15since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 7 first-author · 15 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Overlapping Aware Zone Allocation for LSM Tree-Based Store on ZNS SSDsabstractNVMe Zoned Namespace (ZNS) devices partition the storage space into sequential-write zones, notably reducing the costs of address mapping, garbage collection (GC), and overprovisioning. Log-Structured Merge (LSM) tree-based databases convert random writes into sequential writes and can thus be efficiently handled by ZNS devices. Efficient zone-allocation methods play a pivotal role in maximizing the performance of LSM tree-based store running on ZNS devices. However, existing zone-allocation methods encounter high write-amplification factors due to inaccurate lifetime estimation solely based on the LSM-tree levels. To address this, this paper proposes an overlapping-aware zone-allocation method, termed OAZA, which efficiently selects suitable zones to place data. First, OAZA estimates the data lifetime by considering both the LSM-tree level of the data and the relative data hotness within the same tree level. Secondly, OAZA intelligently selects an appropriate zone to store the data based on the estimated lifetime. Experimental results demonstrate that OAZA outperforms two zone-allocation methods that correlate data lifetime merely to the tree level. Specially, OAZA reduces the amount of GC-induced data copy by average factors of 2.7 × and 1.7× in comparison to the two methods, respectively. Additionally, OAZA achieves an impressively low write-amplification factor of 1.1 ×, outperforming the factors of 1.2× and 1.3× achieved by the two compared methods, respectively. Jingcheng Shen, Linbo Long, Renping Liu 0002, Zhenhua Tan, Congming Gao |
ASPDAC | 4 |
| 2024 | Para-ZNS: Improving Small-Zone ZNS SSDs Parallelism Through Dynamic Zone MappingabstractThe emerging Zoned Namespace (ZNS) interface helps flash-based SSDs achieve high performance by dividing the logical space into fixed-size zones. Typically, a zone is mapped to blocks across multiple dies to achieve I/O parallelism. Small zones can make better use of space and are therefore widely studied. However, a small zone fails to be mapped to blocks residing on all dies, causing underutilized die-level parallelism. Meanwhile, a fine-grained (i.e., plane-level) parallelism is rarely exploited for ZNS SSDs due to a strict limitation mandating that only the same type of operation can be simultaneously performed on the same address across different planes within a die. To address these issues, this paper proposes a novel small-zone ZNS-SSD design with dynamic zone mapping, named Para-ZNS. First, a new parallel block grouping module is devised to group blocks across all planes from multiple dies as a basic unit to be mapped to a zone. Such a basic mapping unit achieves parallelism among multiple dies and plane-level parallelism. Then, a die-parallelism identification module is implemented to locate idle dies. Subsequently, to fully exploit the die-level parallelism, a dynamic zone mapping scheme is employed to intelligently map the basic mapping units on the identified idle dies to open zones. The evaluation results based on a widely-used I/O tester (FIO) demonstrate that Para-ZNS improves the bandwidth by 3.42× on average in comparison to state-of-the-art work. Zhenhua Tan, Linbo Long, Jingcheng Shen, Congming Gao, Renping Liu 0002 |
DATE | 5 |
| 2024 | Hi-ZNS: High Space Efficiency and Zero-Copy LSM-Tree Based Stores on ZNS SSDsabstractThe Zoned Namespace (ZNS) SSD is a newly introduced storage device and provides several new ZNS commands to upper-level applications. Zone-reset command is one of the ZNS commands to erase all the flash blocks within a zone. Since data is grouped and erased in zone units, ZNS SSDs are widely used in LSM-tree-based stores. However, the basic invalidated unit in LSM-tree is an SST/WAL file, which mismatches the erasing unit of a ZNS SSD. Placing different SST/WAL files in the same zone, LSM-tree on ZNS SSDs faces dramatic space amplification and extensive data migration problems. Renping Liu 0002, Peng Chen 0027, Linbo Long, Anping Xiong, Duo Liu 0002 |
ICPP | 1 |
| 2024 | ZNS-Cleaner: Enhancing lifespan by reducing empty erase in ZNS SSDs
Renping Liu 0002, Peng Chen 0027, Linbo Long, Anping Xiong, Duo Liu 0002 |
J. Syst. Archit. | 1 |
| 2024 | WA-Zone: Wear-Aware Zone Management Optimization for LSM-Tree on ZNS SSDsabstractZNS SSDs divide the storage space into sequential-write zones, reducing costs of DRAM utilization, garbage collection, and over-provisioning. The sequential-write feature of zones is well-suited for LSM-based databases, where random writes are organized into sequential writes to improve performance. However, the current compaction mechanism of LSM-tree results in widely varying access frequencies (i.e., hotness) of data and thus incurs an extreme imbalance in the distribution of erasure counts across zones. The imbalance significantly limits the lifetime of SSDs. Moreover, the current zone-reset method involves a large number of unnecessary erase operations on unused blocks, further shortening the SSD lifetime. Considering the access pattern of LSM-tree, this article proposes a wear-aware zone-management technique, termed WA-Zone , to effectively balance inter- and intra-zone wear in ZNS SSDs. In WA-Zone, a wear-aware zone allocator is first proposed to dynamically allocate data with different hotness to zones with corresponding lifetimes, enabling an even distribution of the erasure counts across zones. Then, a partial-erase-based zone-reset method is presented to avoid unnecessary erase operations. Furthermore, because the novel zone-reset method might lead to an unbalanced distribution of erasure counts across blocks in a zone, a wear-aware block allocator is proposed. Experimental results based on the FEMU emulator demonstrate the proposed WA-Zone enhances the ZNS-SSD lifetime by 5.23×, compared with the baseline scheme. Linbo Long, Shuiyong He, Jingcheng Shen, Renping Liu 0002, Zhenhua Tan, Congming Gao, Duo Liu 0002, Kan Zhong |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | Optimizing Garbage Collection for ZNS SSDs via In-storage Data Migration and Address RemappingabstractThe NVMe Zoned Namespace (ZNS) is a high-performance interface for flash-based solid-state drives (SSDs), which divides the logical address space into fixed-size and sequential-write zones. Meanwhile, ZNS SSDs eliminate in-device garbage collection (GC) by shifting the responsibility of GC to the host. However, the host-side GC of ZNS SSDs is not efficient. On the one hand, data migration during GC first moves data to the host buffer and then writes back the transferred data to the new location in the SSD, resulting in an unnecessary end-to-end transfer overhead. On the other hand, due to the pre-configured mapping between zones and blocks, GC incurs a large block-to-block rewrite overhead, i.e., even if most of the data in a block of the victim zone is valid, the valid data will still be rewritten to another block in the target zone. To address these issues, this article proposes a novel ZNS SSD design that features dynamic zone mapping, termed Brick-ZNS . Brick-ZNS implements two key functionalities: in-storage data migration and address remapping. New ZNS commands are first designed to realize in-storage data migration to avoid the end-to-end transfer overhead of GC while ensuring performance predictability. Then, a remapping strategy exploiting parallel physical blocks is proposed to reduce the large block-to-block rewrite overhead while ensuring zone-level access parallelism. The basic idea of the strategy is to directly remap the parallel physical blocks with a sufficient amount of valid data in the victim zone to the target zone, hence avoiding the large block-to-block rewrite overhead. Based on a full-stack SSD emulator, the evaluation results show that Brick-ZNS improves write throughput by 25% and SSD lifetime by 1.41×. Zhenhua Tan, Linbo Long, Jingcheng Shen, Renping Liu 0002, Congming Gao, Kan Zhong |
ACM Trans. Archit. Code Optim. | 4 |
| 2024 | Fair-ZNS: Enhancing Fairness in ZNS SSDs Through Self-Balancing I/O SchedulingabstractThe NVMe Zoned Namespace (ZNS) is a new type of storage interface, which divides logical address space into fixed-size zones, and each zone strictly follows a sequential write constraint with a write pointer. Owing to the sequential write constraint of the ZNS, I/O requests would not be scheduled arbitrarily like the traditional SSDs with block interface. When multiple applications concurrently access one ZNS SSD hardware, the constraint deteriorates I/O blocking and causes huge unfairness. To resolve the problem, we propose a self-balance I/O scheduling dedicated for ZNS SSDs, called Fair-ZNS, to balance the slowdown among multiple applications and ensure fairness. Fair-ZNS identifies the unfair requests by the maximum slowdown value, and violently schedules these requests into the head of the queues overcoming the sequential write constraint. To eliminate the negative effect of the violent scheduling, Fair-ZNS deploys a self-balancing coordinator to fine-tune the order of the requests. Comprehensive evaluations show that Fair-ZNS alleviates I/O blocking and reduces average waiting time by 8.3×, increases fairness by 2.3×, and decreases the max slowdown by 5.1× averagely when compared to the current ZNS SSDs. Renping Liu 0002, Zhenhua Tan, Linbo Long, Duo Liu 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Optimizing Data Migration for Garbage Collection in ZNS SSDsabstractZNS SSDs shift the responsibility of garbage collection (GC) to the host. However, data migration in GC needs to move data to the host's buffer first and write back to the new location, resulting in an unnecessary end-to-end transfer overhead. Moreover, due to the pre-configured mapping between zones and blocks, GC needs to perform a large number of unnecessary block-to-block data migrations between zones. To address these issues, this paper proposes a simple and efficient data migration method, called IS-AR, with in-storage data migration and address remapping. Based on a full-stack SSD emulator, our evaluation shows that IS-AR reduces GC latency by 6.78× and improves SSD lifetime by 1.17× on average. Zhenhua Tan, Linbo Long, Renping Liu 0002, Congming Gao |
DATE | 3 |
| 2023 | ADAR: Application-Specific Data Allocation and Reprogramming Optimization for 3-D TLC Flash MemoryabstractHigh bit-density flash memories, such as triple-level cell (TLC) and quad-level cell (QLC), have been widely used in flash memory-based storage systems, offering significantly high capacity. However, these high bit-density flash memories suffer from asymmetric access performance on the different pages that sharing the same physical cells. Meanwhile, 3-D flash memory adopts stacking technology to increase capacity and reduce cost per bit. The flash unit can be reprogrammed many times as long as the voltage increases. The reprogramming technology is also an effective solution for further increasing the 3-D flash capacity, allowing multiple program operations in an erase cycle. Considering the restrictions of reprogram operations, solid-state drives (SSDs) should capture the access pattern to perform more reprogramming operations to realize the joint optimization of read and write performance. In this work, we propose an application-specific data allocation and reprogramming technique named ADAR to enhance the read and write performance of 3-D TLC flash memory-based SSDs. The core idea is to allocate low-latency least significant bit (LSB) and central significant bit pages to frequently updated write data (termed hot write data) to improve the write performance, and reprogram the pages from high-latency pages (e.g., most significant bit page) to low-latency pages (e.g., LSB page mode) to enhance the read performance while initially storing frequently read data (termed hot read data) in high-latency pages. We explored data access patterns and designed an effective hotness identification method to present a new data allocation and reprogramming technique for 3-D TLC flash memory. Based on a modified 3-D TLC SSD simulator with typical workloads, our evaluation showed that our technique achieved 35.36% and 25.72% performance improvements in read and write latencies, respectively. Linbo Long, Jinpeng Huang, Congming Gao, Duo Liu 0002, Renping Liu 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Optimizing data placement and size configuration for morphable NVM based SPM in embedded multicore systems
Linbo Long, Jinpei Du, Xuxu Deng, Renping Liu 0002, Yan Wang 0022 |
Future Gener. Comput. Syst. | 4 |
| 2022 | Self-Adapting Channel Allocation for Multiple Tenants Sharing SSD DevicesabstractSolid-state drives (SSDs) have been widely deployed in high-performance data center environments, where multiple tenants usually share the same hardware. However, traditional SSDs distribute the users’ incoming data uniformly across all SSD channels, which leads to numerous access conflicts. Meanwhile, SSDs that blindly allocate one or several channels to one tenant sacrifice device parallelism and capacity. When SSDs are shared by tenants with different access patterns, inappropriate channel allocation results in SSD performance degradation. In this article, we propose a self-adapting channel allocation mechanism, named SSDKeeper, for multiple tenants that share one SSD. SSDKeeper employs a machine learning-assisted algorithm to take full advantage of SSD parallelism while providing performance isolation. By collecting multitenant access patterns, SSDKeeper predicts an optimal channel allocation strategy for multiple tenants using the well-trained model. To further consume the blocks in different channels evenly, SSDKeeper equips with a novel channel swap scheme to prolong the SSD lifespan. Comparing with traditional SSDs, SSDKeeper reduces the overall latency of read and write by 12.6% and the lifespan is prolonged up to$3.7\times $. Renping Liu 0002, Duo Liu 0002, Xianzhang Chen, Yujuan Tan, Runyu Zhang 0002, Liang Liang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Improving Fairness for SSD Devices through DRAM Over-Provisioning Cache ManagementabstractModern NVMe SSDs have been widely deployed in multi-tenant cloud computing environments or multi-programming systems. When multiple applications concurrently access one SSD hardware, unfairness within the shared SSD will slow down the application significantly and lead to a violation of service level objectives. However, traditional data cache management within SSDs mainly focuses on improving cache hit ratio, which causes data cache contention and sacrifices fairness among multiple applications. In this paper, we propose a DRAM-based Over-Provisioning (OP) cache management mechanism, named Justitia, to reduce data cache contention and improve fairness for modern SSDs. Justitia consists of two stages includingStatic-OPstage andDynamic-OPstage. Through the novel OP mechanism in the two stages, Justitia reduces the max slowdown by$4.5\times$on average. At the same time, Justitia increases fairness by$20.6\times$and buffer hit ratio by$19.6\%$averagely, compared with the traditional shared mechanism. Renping Liu 0002, Zhenhua Tan, Linbo Long, Yu Wu 0016, Yujuan Tan, Duo Liu 0002 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Forseti: An Efficient Basic-block-level Sensitivity Analysis Framework Towards Multi-bit FaultsabstractThe per-instruction sensitivity analysis framework is developed to evaluate the resiliency of a program and identify the segments of the program needing protection. However, for multi-bit hardware faults, the per-instruction sensitivity analysis frameworks can cause large overhead for redundant analyses. In this paper, we propose a basic-block-level sensitivity analysis framework, Forseti, to reduce the analysis overhead in analyzing impacts of modern microprocessors' multi-bit faults on programs. We implement Forseti in LLVM and evaluate it with five typical workloads. Extensive experimental results show that Forseti can achieve more than 90% sensitivity classification accuracy and 6.16× speedup over instruction-level analysis. Jinting Ren, Xianzhang Chen, Duo Liu 0002, Moming Duan, Renping Liu 0002, Chengliang Wang 0002 |
DATE | 5 |
| 2021 | MobileRE: A replicas prioritized hybrid fault tolerance strategy for mobile distributed system
Yu Wu 0016, Duo Liu 0002, Xianzhang Chen, Jinting Ren, Renping Liu 0002, Yujuan Tan, Ziling Zhang |
J. Syst. Archit. | 5 |
| 2021 | Self-Balancing Federated Learning With Global Imbalanced Data in Mobile SystemsabstractFederated learning (FL) is a distributed deep learning method that enables multiple participants, such as mobile and IoT devices, to contribute a neural network while their private training data remains in local devices. This distributed approach is promising in the mobile systems where have a large corpus of decentralized data and require high privacy. However, unlike the common datasets, the data distribution of the mobile systems is imbalanced which will increase the bias of model. In this article, we demonstrate that the imbalanced distributed training data will cause an accuracy degradation of FL applications. To counter this problem, we build a self-balancing FL framework named Astraea, which alleviates the imbalances by 1) Z-score-based data augmentation, and 2) Mediator-based multi-client rescheduling. The proposed framework relieves global imbalance by adaptive data augmentation and downsampling, and for averaging the local imbalance, it creates the mediator to reschedule the training of clients based on Kullback-Leibler divergence (KLD) of their data distribution. Compared with FedAvg, the vanilla FL algorithm, Astraea shows +4.39 and +6.51 percent improvement of top-1 accuracy on the imbalanced EMNIST and imbalanced CINIC-10 datasets, respectively. Meanwhile, the communication traffic of Astraea is reduced by 75 percent compared to FedAvg. Moming Duan, Duo Liu 0002, Xianzhang Chen, Renping Liu 0002, Yujuan Tan, Liang Liang 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | SSDKeeper: Self-Adapting Channel Allocation to Improve the Performance of SSD DevicesabstractSolid state drives (SSDs) have been widely deployed in high performance data center environments, where multiple tenants usually share the same hardware. However, traditional SSDs distribute the users' incoming data uniformly across all SSD channels, which leads to numerous access conflicts. Meanwhile, SSDs that statically allocate one or several channels to one tenant sacrifice device parallelism and capacity. When SSDs are shared by tenants with different access patterns, inappropriate channel allocation results in SSDs performance degradation. In this paper, we propose a self-adapting channel allocation mechanism, named SSDKeeper, for multiple tenants to share one SSD. SSDKeeper employs a machine learning assisted algorithm to take full advantage of SSD parallelism while providing performance isolation. By collecting multi-tenant access patterns and training a model, SSDKeeper selects an optimal channel allocation strategy for multiple tenants with the lowest overall response latency. Experimental results show that SSDKeeper improves the overall performance by 24% with negligible overhead. Renping Liu 0002, Xianzhang Chen, Yujuan Tan, Runyu Zhang 0002, Liang Liang 0002, Duo Liu 0002 |
IPDPS | 1 |
| 2020 | Separable Binary Convolutional Neural Network on Embedded SystemsabstractWe have witnessed the tremendous success of deep neural networks. However, this success comes with the considerable memory and computational costs which make it difficult to deploy these networks directly on resource-constrained embedded systems. To address this problem, we propose TaijiNet, a separable binary network, to reduce the storage and computational overhead while maintaining a comparable accuracy. Furthermore, we also introduce a strategy called partial binarized convolution which binarizes only unimportant kernels to efficiently balance network performance and accuracy. Our approach is evaluated on the CIFAR-10 and ImageNet datasets. The experimental results show that with the proposed TaijiNet, the separable binary versions of AlexNet and ResNet-18 can achieve 26× and 6.4× compression rates with comparable accuracy when comparing with the full-precision versions respectively. In addition, by adjusting the PCA threshold, the xnor version of Taiji-AlexNet improves accuracy by 4-8 percent comparing with other state-of-the-art methods. Renping Liu 0002, Xianzhang Chen, Duo Liu 0002, Yingjian Ling, Weilue Wang, Yujuan Tan, Chunhua Xiao, Chaoshu Yang, Runyu Zhang 0002, Liang Liang 0002 |
IEEE Trans. Computers | 1 |
| 2020 | Downsizing Without Downgrading: Approximated Dynamic Time Warping on Nonvolatile MemoriesabstractIn recent years, time-series data have emerged in a variety of application domains, such as wireless sensor networks and surveillance systems. To identify the similarity between time-series data, the Euclidean distance and its variations are common metrics that quantify the differences between time-series data. However, the Euclidean distance is limited by its inability to elastically shift with the time axis, which motivates the development of dynamic time warping (DTW) algorithms. While DTW algorithms have been proven very useful in diversified applications like speech recognition, their efficacy might be seriously affected by the resolution of the time-series data. However, high-resolution time-series data might take up a gigantic amount of main memory and storage space, which will slow down the DTW analysis procedure. This makes the upscaling of DTW analysis more challenging, especially for in-memory data analytics platforms with limited nonvolatile memory space. In this paper, we propose a strategy to downsample time-series data to significantly reduce their size without seriously affecting the precision of the results obtained by DTW algorithms (downsizing without downgrading). In other words, this paper proposes a technique to remove the unimportant details that are largely ignored by DTW algorithms. The efficacy of the proposed technique is verified by a series of experimental studies, where the results are quite encouraging. Duo Liu 0002, Xingni Li, Po-Chun Huang, Yingjian Ling, Kan Zhong, Renping Liu 0002, Xianzhang Chen, Liang Liang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2019 | FitCNN: A cloud-assisted and low-cost framework for updating CNNs on IoT devices
Duo Liu 0002, Chaoshu Yang, Xianzhang Chen, Jinting Ren, Renping Liu 0002, Moming Duan, Yujuan Tan, Liang Liang 0002 |
Future Gener. Comput. Syst. | 6 |
| 2019 | Towards Fast and Lightweight Checkpointing for Mobile Virtualization Using NVRAMabstractCheckpointing is a key enabler of hibernation, live migration and fault-tolerance for virtual machines (VMs) in mobile devices. However, checkpointing a VM is usually heavyweight: the VM's entire memory needs to be dumped to storage, which induces a significant amount of (slow) I/O operations, degrading system performance and user experience. In this paper, we propose FLIC, a fast and lightweight checkpointing machinery for virtualized mobile devices by taking advantages of recent byte-addressable, non-volatile memory (NVRAM). Instead of saving the VM's entire memory to storage, we store its working set pages in NVRAM, avoiding accessing slow flash memory (compared to server-grade SSDs). To further reduce the write activities to flash memory, we propose an energy-efficient data deduplication to eliminate redundant data in VM snapshot and save storage space. Experimental results based on an Exynos 5250 SoC show that our approach can effectively improve the performance of checkpointing in mobile virutalization and save energy. Kan Zhong, Duo Liu 0002, Yunsong Wu, Linbo Long, Weichen Liu 0001, Jinting Ren, Renping Liu 0002, Liang Liang 0002, Zili Shao, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 7 |