VLDB 2026 Research / reviewers in the wild / expert
Jiapin Wang
dblp:231/4026
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2026
0009-0007-8032-0862ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CEMU: Enabling Full-System Emulation of Computational Storage Beyond Hardware LimitsabstractComputational storage drives (CSDs) present a promising approach to improve system performance through near data processing in SSDs. However, current research platforms are fragmented and inadequate to explore the full design space of CSD systems. Existing hardware and emulator platforms are constrained by physical compute resources, while simulators lack full-system fidelity. To address the problems, we introduce CEMU, a new software-based CSD emulation platform that enables full-system research. It consists of a CSD device emulator and a CSD-oriented software stack. Through a novel virtual machine freezing mechanism, CSD emulation achieves high configurability. While the CSD can utilize the host CPU to physically perform computation to preserve full-system behaviors, the computational delay can be modeled separately to emulate CSDs with CPU-unbounded high computing power. The software stack is designed with two principles, adhering to recent industry CSD standards and being compatible with the existing I/O stack, which is achieved via a newly developed file system FDMFS. We verify CEMU's emulation fidelity across a range of applications by benchmarking against actual CSD hardware, demonstrating average end-to-end performance accuracy of 95% or higher. We also use two case studies on large language model training and LevelDB to demonstrate that CEMU is effective in exploring CSD system research and can uncover insights that have not been discovered in previous research platforms. Jiapin Wang, You Zhou 0009, Kai Lu 0002, Jiguang Wan 0001, Fei Wu 0005, Tao Lu 0014 |
ASPLOS (2) | 2 |
| 2026 | ASIC-based Compression Accelerators for Storage Systems: Design, Placement, and Profiling Insights
Tao Lu 0014, Jiapin Wang, Yelin Shan, Xiang Chen 0028 |
EuroSys | 2 |
| 2025 | StreamCSD: SSD-Autonomous Stream Management via In-Storage Content LearningabstractWrite amplification (WA) from migrating valid pages during garbage collection (GC) degrades SSD performance and lifespan. Although stream management based on high-level software semantics reduces WA, existing solutions require host modifications, hindering their adoption. We introduce StreamCSD, an SSD-autonomous stream management approach using in-storage content learning, eliminating host-side changes. Leveraging compression ratios from embedded compressors in computational storage drives (CSDs), StreamCSD employs a streaming Kmeans algorithm to cost-efficiently cluster data into streams. Evaluations show that StreamCSD reduces WA from 1.7 to 1.06 under multimodal generative AI workloads, matching state-of-the-art methods with minimal impact on bandwidth. StreamCSD operates without host modifications, promoting broader adoption of multi-stream SSDs. Xiang Chen 0028, Yelin Shan, Jiapin Wang, Yunxin Huang, Yafei Yang, Tao Lu 0014, You Zhou 0009, Fei Wu 0005 |
DAC | 4 |
| 2024 | HA-CSD: Host and SSD Coordinated Compression for Capacity and PerformanceabstractIntegrating data compression capability into SSDs has demonstrated great potential to improve the utilization and lifetime of the storage device and also the performance of the entire system. It is advocated to add a hardware engine into the SSD for low-latency compression and decompression. However, this requires a new and long hardware product development cycle, which would prevent current storage systems from reaping the benefits of in-SSD compression. In this paper, we explore a software-based in-SSD compression solution, which can be delivered to users quickly through a simple SSD firmware update. The most critical challenge is the severe performance bottleneck caused by compression and decompression, as the in-SSD embedded CPU has quite limited computing power. To tackle this challenge, we propose a host-assisted computational storage device, called HA-CSD. It employs an offline, data hotness- and compressibility-aware compression strategy to remove compression from the critical write I/O path. A novel decompression architecture is devised to utilize the powerful host CPU for fast decompression. We implement HA-CSD in a commercial enterprise SSD with a code change of more than 25K lines in the host NVMe driver and SSD firmware. Experimental results show that HA-CSD achieves 2.1GB/s and 5.2GB/s read and write bandwidth. Compared with RocksDB built-in compression, HA-CSD can increase the YCSB benchmark throughput by up to 5.7×, and improve the host CPU efficiency significantly. Xiang Chen 0028, Tao Lu 0014, Jiapin Wang, Guangchun Xie, Xueming Cao, Yuanpeng Ma, Bing Si, Yunxin Huang, Yafei Yang, You Zhou 0009, Fei Wu 0005 |
IPDPS | 3 |
| 2023 | Optimizing the Performance of NDP Operations by Retrieving File Semantics in StorageabstractIn-storage Near-Data Processing (NDP) architectures can reduce data movement between the host and the storage device by offloading computing tasks to the storage. This encourages many studies on building NDP applications, such as recommendation systems and databases, on computational SSDs. However, in the data path of existing NDP architectures, an NDP application has to find out the address of the requested file data by calling the I/O stacks of the kernel on the host, which incurs large overhead for transferring data between the host and the computational SSD. In this paper, we present File Semantics Retriever (FSR) to optimize the data path of NDP architectures by locating and fetching the requested file data directly in the computational SSD. The key idea is to recognize the file system layout and the metadata structures in the storage with the collaboration of a user-space library and a handler in the firmware of the computational SSD. We implement a prototype of FSR and evaluate it on the Cosmos plus OpenSSD, a widely-used computational SSD platform. The experimental results show that FSR outperforms existing NDP architectures in both benchmarks and real-world NDP applications. Xianzhang Chen, Jiapin Wang, Duo Liu 0002, Yujuan Tan, Ao Ren |
DAC | 4 |
| 2023 | FSR: A host-storage collaborative mechanism for data path optimization of NDP operations
Qiao Sun 0007, Xianzhang Chen, Jiapin Wang, Shukan Liu |
J. Syst. Archit. | 4 |
| 2022 | eRDAC: Efficient and Reliable Remote Direct Access and Control for Embedded SystemsabstractEmerging embedded systems, such as autonomous vehicles, demand highly efficient remote data transfer, whereas existing networking hardware and protocols cause high communication latency and CPU consumption. In this article, we propose embedded RDAC (eRDAC), an efficient and reliable remote direct access and control solution for embedded systems. The proposed remote access controller in eRDAC has a two-layer protocol offload engine that employs the command/response protocol on UDP to ensure the data reliability and security, and a multichannel DMA controller with configurable priority to improve the efficiency. Besides, a reusable hardware Ethernet MAC is implemented to support not only remote access commands but also standard Ethernet communication. We implement eRDAC on FPGA and the corresponding software in the Linux system. Experimental results show that eRDAC can reduce the latency of remote I/O reading/writing by 74.3%/74.9% ($3.76\times /3.98\times $performance improvement) and reduce the latency of remote memory reading and writing with 1024B by 54.2% compared to the socket-based communication. Meanwhile, eRDAC can cut off the consumption of the remote processor and achieve 0.250mJ/Mb energy consumption with only 25-mW power. Xianzhang Chen, Duo Liu 0002, Weigong Zhang, Jiapin Wang, Rongwei Zheng, Yujuan Tan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Horae: A Hybrid I/O Request Scheduling Technique for Near-Data Processing-Based SSDabstractNear-data processing (NDP) architecture is promised to break the bottleneck of data movement in many scenarios (e.g., databases and recommendation systems), which limits the efficiency of data processing. Different from traditional SSD, NDP-based SSD not only needs to handle normal I/Os (e.g., read and write), but also needs to handle NDP requests that contain data processing operations. NDP and normal I/O requests share some function units of NDP-based SSD, such as flash chips and embedded processors. However, existing works ignore the resource competition between normal I/Os and NDP requests, which drastically degrades the performance. In this article, we propose a novel scheduling technique called Horae, which can efficiently schedule hybrid NDP-normal I/O requests in NDP-based SSD to improve performance. Horae exploits the critical paths on critical resources to maximize the parallelism of multiple stages of requests. The experimental results on typical workloads show that Horae can significantly improve the performance of hybrid NDP-normal I/O requests over the state-of-the-art scheduling algorithms of NDP-based SSDs. Xianzhang Chen, Duo Liu 0002, Jiapin Wang, Zhaoyang Zeng, Yujuan Tan, Lei Qiao 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | SENTunnel: Fast Path for Sensor Data Access on Automotive Embedded Systems
Rongwei Zheng, Xianzhang Chen, Duo Liu 0002, Jiapin Wang, Ao Ren, Chengliang Wang 0002, Yujuan Tan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |