EDBT 2026 Demo / reviewers in the wild / expert
Zelin Du
dblp:337/2093
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Expanding Logical Space Freely: A Memory-efficient Mapping Table Design for Compressional SSDsabstractCompressional SSDs can be a mixed blessing. While they offer users expanded logical space beyond the physical capacity, they complicate the Flash Translation Layer (FTL) design by requiring a larger Logical-to-Physical (L2P) address mapping, which places a heavier burden on the limited in-SSD memory, leading to a degraded I/O performance. In this paper, we aim to reduce the memory footprint of the L2P mapping table in compressional SSDs by proposing a novel N-to-1 L2P mapping table design that consolidates multiple logical entries into a single entry. This approach eliminates the duplication of physical page numbers when a physical page contains several compressed logical pages. To accommodate the dynamic compression ratios of real-world workloads, we introduce promotion and demotion that enable mapping table entries to migrate between pages with different compression ratios. Additionally, to address the issue of partial invalidation-where some compressed logical pages within a physical page are invalid due to the N-to-1 mapping-we present a compression-aware garbage collection algorithm aimed at minimizing the number of copy operations for partially invalid physical pages. We have implemented our design in MQsim, a widely used SSD simulator, and have conducted a series of experiments to evaluate the effectiveness of the proposed techniques. The results demonstrate that our approach significantly reduces the mapping table size in compressional SSDs, leading to an improved mapping table cache hit ratio and a reduced I/O latency compared to traditional compressional SSDs. Zixuan Huang 0011, Tianyu Wang 0009, Kecheng Huang, Zelin Du, Zili Shao |
DAC | 4 |
| 2025 | A Practical Learning-Based FTL for Memory Constrained Mobile Flash StorageabstractThe rapidly growing mobile market is pushing flash storage manufacturers to expand capacity into the terabyte range. However, this presents a significant challenge for mobile storage management: more logical-to-physical page mappings are desired to be efficiently managed and cached while the available caching space is extremely limited. This motivates us to shift toward a new learning-based paradigm: rather than maintaining mappings for individual pages, the learning-based approach can represent mapping relationships for a set of continuous pages. However, to construct linear models, existing methods that either consume the already-limited memory space or reuse flash garbage collection demonstrate poor model construction capabilities or significantly degrade flash performance, making them impractical for real-world use. In this paper, we propose LFTL, a practical, learning-based on-demand flash translation layer design for flash management in mobile devices. In contrast to prior work that centered around gathering sufficient mappings for linear model construction, our key insight is that linear patterns can be extracted and refined by leveraging the orderly, LPA-aligned write stream typical of mobile devices. By doing this, highly accurate linear models can be constructed regardless of the constraints of mobile device's cache limitation. We have implemented a fully functional prototype of LFTL based on FEMU. Our evaluation results show that LFTL is more adaptable to memory-constrained storage devices than state-of-the-art learning-based approaches. Zelin Du, Kecheng Huang, Tianyu Wang 0009, Xin Yao 0008, Renhai Chen, Zili Shao |
DATE | 1 |
| 2024 | PipeSSD: A Lock-free Pipelined SSD Firmware Design for Multi-core ArchitectureabstractModern SSD firmware is continuously optimized for higher parallelism to match the growing frontend PCIe bandwidth with more backend flash channels. Although a multi-core microprocessor is typically adopted to concurrently process independent NVMe requests from multiple NVMe queues, the existing one-to-many thread-request mapping model with each thread serving one or more incoming I/O requests has poor scalability due to severe lock contention problem, especially in cache management. Zelin Du, Shaoqi Li, Zixuan Huang 0011, Jin Xue, Kecheng Huang, Tianyu Wang 0009, Zili Shao |
DAC | 1 |
| 2023 | Accelerating DNN Inference with Heterogeneous Multi-DPU EnginesabstractThe Deep Learning Processor (DPU) programmable engine released by the official Xilinx Vitis AI toolchain has become one of the commercial off-the-shelf (COTS) solutions for Convolutional Neural Networks (CNNs) inference on Xilinx FPGAs. While modern FPGA devices generally have enough hardware resources to accommodate multi-DPUs simultaneously, the Xilinx toolchain currently only supports the deployment of multiple homogeneous DPUs engines that running independent inference tasks (task-level parallelism). In this work, we demonstrate that deployment of multiple heterogeneous DPU engines makes better resource efficiency for a given FPGA device. Moreover, we show that pipelined execution of a CNN inference task over heterogeneous multi-DPU engines may further improve overall inference throughput with carefully designed CNN layers-to-DPU mapping and scheduling. Finally, for a given CNN model and an FPGA device, we propose a comprehensive framework that automatically determines the optimal heterogeneous DPU deployment, and adaptively chooses the execution scheme between task-level and pipelined parallelism. Compared with the state-of-the-art solution with homogeneous multi-DPU engines and network-level parallelism, the proposed framework shows an average improvement of 13% (up-to 19%) and 6.6% (up-to 10%) on the Xilinx Zynq UltraScale+ MPSoC ZCU104 and ZCU102 platforms, respectively. Zelin Du, Wei Zhang 0173, Zimeng Zhou, Zili Shao, Lei Ju 0001 |
DAC | 1 |
| 2023 | Lightning Talk: Model, Framework and Integration for In-Storage Computing with Computational SSDsabstractIn-storage computing with computational SSDs is emerging as one effective solution for I/O bottlenecks in big data applications such as AI learning model training. Specifically, with in-SSD computing, computation can be pushed down to SSDs and the volume of the output data that will be transferred back to the host can be greatly reduced. However, there are several fundamental issues for applications to fully exploit in-SSD computing with simple and efficient function offloading. In this paper, we present three challenges for in-SSD computing, namely, data model, programming framework, and storage/computing integration, and discuss possible research directions. Tianyu Wang 0009, Jin Xue, Zelin Du, Yaotian Cui, Zili Shao |
DAC | 3 |
| 2023 | A Comprehensive Memory Management Framework for CPU-FPGA Heterogenous SoCsabstractEfficient utilization of restrained memory resources is of paramount importance in CPU-FPGA heterogeneous multiprocessor system-on-chip (HMPSoC)-based system design for memory-intensive applications. State-of-the-art high level synthesis (HLS) tools rely on the system programmers to manually determine the data placement within the complex memory hierarchy. Different data placement policies may lead to different system performance, and finding an optimal data placement policy is a nontrivial problem. For instance, we show counter-intuitive results that traditional frequency and locality-based data placement strategy designed for CPU architecture leads to nonoptimal system performance in CPU-FPGA HMPSoCs. In this work, we first propose an automatic data placement framework for field programmable gate array (FPGA) kernels to determine whether each array object should be accessed via the on-chip BRAM, shared CPU L2-cache, or DDR memory to achieve the optimal performance. Moreover, we find that when the CPU kernel and the FPGA kernel are executed in parallel, memory contentions may degrade the performance and the optimal data placement policy designed for the FPGA kernel alone will not achieve the optimal overall system performance. In this article, we proposed to use cache partitioning to alleviate the impact brought by memory contentions. We extend the framework designed for FPGA by adding the cross-layer memory contentions analysis to automatically generate an optimal data placement policy and cache partitioning mechanism for the parallel executing kernels. The proposed data placement framework can be seamlessly integrated with the commercial Vivado HLS. The experimental results on the Zedboard platform show an average$1.5\times $performance speedup for FPGA kernels compared with a greedy-based allocation strategy. When FPGA kernels and CPU kernels are executed in parallel, the FPGA kernel and the CPU kernel have a performance speedup of$1.62\times $and$1.10\times $on average, respectively. Zelin Du, Qianling Zhang, Mao Lin, Shiqing Li, Xin Li 0137, Lei Ju 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |