VLDB 2026 Research / reviewers in the wild / expert
Xiang Chen 0028
dblp:64/3062-28
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-6178-6217ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ASIC-based Compression Accelerators for Storage Systems: Design, Placement, and Profiling Insights
Tao Lu 0014, Jiapin Wang, Yelin Shan, Xiang Chen 0028 |
EuroSys | 5 |
| 2026 | Crash-Consistent SSD Array With Hardware-Guaranteed Transactional AtomicityabstractCrash consistency is a critical challenge that storage systems need to address carefully in multiple layers, including the database, file system, and block-layer RAID. Software approaches typically resort to logging to ensure transactional atomicity, which causes significant performance overhead and write amplification over underlying high-speed SSDs. Motivated by the out-of-place update feature of SSD-internalflash translation layer (FTL), previous studies propose crash-consistent SSDs to offload transactional atomicity guarantee and demonstrate their effectiveness in eliminating software-based logging overheads. However, existing hardware offloading approaches only consider single-SSD systems and would fail in an SSD array. This paper presents a crash-consistent SSD array, using FTLs and coordinating multiple SSDs to provide transactional atomicity across the array. The key is to design anarray-wide transaction commit (ARC)protocol, which resolves the multi-SSD coordination challenge and tolerates disk failures. We implement an ARC array manager and ARC SSDs to verify the design. Two case studies are conducted, where the ARC array is utilized to address the transaction logging overhead in the SQLite database and the stripe write-hole problem in the RAID subsystem. Experimental results demonstrate that the ARC SSD array can improve system performance by 32% to 93% and reduce write amplification by 39% to 49%, on average, compared with software-based logging approaches. Zeyu Niu, Xiang Chen 0028, You Zhou 0009, Zibin Sun, Zhihu Tan, Changsheng Xie 0001, Fei Wu 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | Computational Burst Buffers: Accelerating HPC I/O via In-Storage Compression OffloadingabstractBurst buffers (BBs) act as an intermediate storage layer between compute nodes and parallel file systems (PFS), effectively alleviating the I/O performance gap in high-performance computing (HPC). As scientific simulations and AI workloads generate larger checkpoints and analysis outputs, BB capacity shortages and PFS bandwidth bottlenecks are emerging, and CPU-based compression is not an effective solution due to its high overhead. We introduceComputational Burst Buffers(CBBs), a storage paradigm that embeds hardware compression engines such as application-specific integrated circuit (ASIC) inside computational storage drives (CSDs) at the BB tier. CBB transparently offloads both lossless and error-bounded lossy compression from CPUs to CSDs, thereby (i) expanding effective SSD-backed BB capacity, (ii) reducing BB–PFS traffic, and (iii) eliminating contention and energy overheads of CPU-based compression. Unlike prior CSD-based compression designs targeting databases or flash caching, CBB co-designs the burst-buffer layer and CSD hardware for HPC and quantitatively evaluates compression offload in BB–PFS hierarchies. We prototype CBB using a PCIe 5.0 CSD with an ASIC Zstd-like compressor and an FPGA prototype of an SZ entropy encoder, and evaluate CBB on a 16-node cluster. Experiments with four representative HPC applications and a large-scale workflow simulator show up to 61% lower application runtime, 8–12× higher cache hit ratios, and substantially reduced compute-node CPU utilization compared to software compression and conventional BBs. These results demonstrate that compression-aware BBs with CSDs provide a practical, scalable path to next-generation HPC storage. Xiang Chen 0028, Bing Lu 0001, Haoquan Long, Huizhang Luo, Yili Ma, Guangming Tan, Dingwen Tao, Fei Wu 0005, Tao Lu 0014 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | StreamCSD: SSD-Autonomous Stream Management via In-Storage Content LearningabstractWrite amplification (WA) from migrating valid pages during garbage collection (GC) degrades SSD performance and lifespan. Although stream management based on high-level software semantics reduces WA, existing solutions require host modifications, hindering their adoption. We introduce StreamCSD, an SSD-autonomous stream management approach using in-storage content learning, eliminating host-side changes. Leveraging compression ratios from embedded compressors in computational storage drives (CSDs), StreamCSD employs a streaming Kmeans algorithm to cost-efficiently cluster data into streams. Evaluations show that StreamCSD reduces WA from 1.7 to 1.06 under multimodal generative AI workloads, matching state-of-the-art methods with minimal impact on bandwidth. StreamCSD operates without host modifications, promoting broader adoption of multi-stream SSDs. Xiang Chen 0028, Yelin Shan, Jiapin Wang, Yunxin Huang, Yafei Yang, Tao Lu 0014, You Zhou 0009, Fei Wu 0005 |
DAC | 2 |
| 2024 | HA-CSD: Host and SSD Coordinated Compression for Capacity and PerformanceabstractIntegrating data compression capability into SSDs has demonstrated great potential to improve the utilization and lifetime of the storage device and also the performance of the entire system. It is advocated to add a hardware engine into the SSD for low-latency compression and decompression. However, this requires a new and long hardware product development cycle, which would prevent current storage systems from reaping the benefits of in-SSD compression. In this paper, we explore a software-based in-SSD compression solution, which can be delivered to users quickly through a simple SSD firmware update. The most critical challenge is the severe performance bottleneck caused by compression and decompression, as the in-SSD embedded CPU has quite limited computing power. To tackle this challenge, we propose a host-assisted computational storage device, called HA-CSD. It employs an offline, data hotness- and compressibility-aware compression strategy to remove compression from the critical write I/O path. A novel decompression architecture is devised to utilize the powerful host CPU for fast decompression. We implement HA-CSD in a commercial enterprise SSD with a code change of more than 25K lines in the host NVMe driver and SSD firmware. Experimental results show that HA-CSD achieves 2.1GB/s and 5.2GB/s read and write bandwidth. Compared with RocksDB built-in compression, HA-CSD can increase the YCSB benchmark throughput by up to 5.7×, and improve the host CPU efficiency significantly. Xiang Chen 0028, Tao Lu 0014, Jiapin Wang, Guangchun Xie, Xueming Cao, Yuanpeng Ma, Bing Si, Yunxin Huang, Yafei Yang, You Zhou 0009, Fei Wu 0005 |
IPDPS | 1 |
| 2024 | A Unified Computational Storage and Memory Architecture in the CXL EraabstractComputational Storage Drives (CSDs) integrate compute engines directly within SSDs for efficient near-data processing. With the introduction of Compute Express Link (CXL), memory expanders can similarly become Computational Memory (CM) by offloading computations. However, the integration of specific hardware accelerators in CSDs has posed substantial mass production challenges, a pitfall we anticipate will also affect CM. To address this, we propose an innovative, decoupled architecture that uses CXL switches to separate accelerators from storage and memory devices. Our analysis suggests this architecture effectively sidesteps the anti-mass production issues faced by CSD and CM with an affordable (e.g.,$< 20\%$) performance degradation. Xiang Chen 0028, Fei Wu 0005, Tao Lu 0014 |
NAS | 1 |
| 2023 | ADT-FSE: A New Encoder for SZabstractSZ is a lossy floating-point data compressor that excels in compression ratio and throughput for high-performance computing (HPC), time series databases, and deep learning applications. However, SZ performs poorly for small chunks and has slow decompression. We pinpoint the Huffman tree in the quantization factor encoder as the bottleneck of SZ. In this paper, we propose ADT-FSE, a new quantization factor encoder for SZ. Based on the Gaussian distribution of quantization factors, we design an adaptive data transcoding (ADT) scheme to map quantization factors to codes for better compressibility, and then use finite state entropy (FSE) to compress the codes. Experiments show that ADT-FSE improves the quantization factor compression ratio, compression and decompression throughput by up to 5×, 2× and 8×, respectively, over the original SZ Huffman encoder. On average, SZ_ADT is over 2× faster than ZFP in decompression. Case studies of the TDengine time series database and HDF5 file store confirm that SZ_ADT significantly boosts user-perceived application performance. In addition, ADT-FSE makes the compression ratio prediction of SZ_ADT easy and accurate, and has the potential to dramatically reduce the area size of SZ hardware implementation. Tao Lu 0014, Zibin Sun, Xiang Chen 0028, You Zhou 0009, Fei Wu 0005, Yunxin Huang, Yafei Yang |
SC | 4 |
| 2022 | Error Generation for 3D NAND Flash MemoryabstractThree-dimension (3D) NAND flash memory is the preferred storage component of solid-state drive (SSD) for its high ratio of capacity and cost. Optimizing the reliability of modern SSD needs to test and collect a large amount of real-world error data from 3D NAND flash memory. However, the test costs have surged dozens of times as its capacity increases. It's imperative to reduce the costs of testing denser and high-capacity flash memory. To facilitate it, in this paper, we aim to enable reproducing error data efficiently for 3D NAND flash memory. We use a conditional generative adversarial network (cGAN) to learn the error distribution with multiple interferences and generate diverse error data comparable to the real-world. Evaluation results demonstrate it is feasible and efficient for error generation with cGAN. Fei Wu 0005, Songmiao Meng, Xiang Chen 0028, Changsheng Xie 0001 |
DATE | 4 |
| 2022 | Characterization Summary of Performance, Reliability, and Threshold Voltage Distribution of 3D Charge-Trap NAND Flash MemoryabstractSolid-state drive (SSD) gradually dominates in the high-performance storage scenarios. Three-dimension (3D) NAND flash memory owning high-storage capacity is becoming a mainstream storage component of SSD. However, the interferences of the new 3D charge-trap (CT) NAND flash are getting unprecedentedly complicated, yielding to many problems regarding reliability and performance. Alleviating these problems needs to understand the characteristics of 3D CT NAND flash memory deeply. To facilitate such understanding, in this article, we delve into characterizing the performance, reliability, and threshold voltage ( V th ) distribution of 3D CT NAND flash memory. We make a summary of these characteristics with multiple interferences and variations and give several new insights and a characterization methodology. Especially, we characterize the skewed ( V th ) distribution, ( V th ) shift laws, and the exclusive layer variation in 3D NAND flash memory. The characterization is the backbone of designing more reliable and efficient flash-based storage solutions. Fei Wu 0005, Xiang Chen 0028, Meng Zhang 0014, Yu Wang 0168, Xiangfeng Lu, Changsheng Xie 0001 |
ACM Trans. Storage | 3 |
| 2010 | Cache Blocks: An Efficient Scheme for Solid State Drives without DRAM CacheabstractMost solid state drives use DRAM for device's cache, the volatile memory provides the I/O Caching ability, and maintains the drives' mapping table (indicates the correspondence between physical unit and logical unit). However, when the drives' power shut down unexpected, the volatile DRAM memory may lose the caching data, which did not have time to write to the drives' storage media, so the dirty data generated. This paper proposes an efficient management scheme for low cost Solid State Drives, with low cost ASIC controller chip, only has internal SRAM memory, and no external DRAM. We use a kind of Cache Blocks: when write requests come, write in these areas first, and write in the sequentially physical place, and the limited internal SRAM for the mapping tables maintaining and data transferring. We propose some efficient methods: 1)using flash memory as cache, 2) page mapping for cache blocks regions, block mapping for data blocks regions, 3) binding two planes operation, 4) using the idle internal plane SRAM as data buffer to improve the I/O performance, without DRAM. So we avoid the dirty data when the power loses unexpected. And this scheme is energy-efficient and low cost. We test the scheme on our own SSD test board, with the pool SRAM size, the I/O performances do not decrease two much, and the random write even better about 20%, compare to the SSD with DRAM. The experiment also shows, this scheme cuts about 21% energy than the DRAM architecture. And it may be adapted in consumer electronics area. Fei Wu 0005, Xiang Chen 0028, Jiguang Wan 0001 |
NAS | 2 |