VLDB 2026 Research / reviewers in the wild / expert
Shucheng Wang
dblp:220/0137
· DBLP profile ↗
19ranked-venue papers
10as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 10 first-author · 13 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ZTree: Towards an efficient B+-tree on zoned namespace SSDs
Kaiye Zhou, Shucheng Wang |
Future Gener. Comput. Syst. | 3 |
| 2026 | The Design of Trillion-scale SSD-based Indexing with Deterministic Latency for Cloud Block StorageabstractCloud block storage (CBS) provides virtual disks with block-level accessibility. The petabyte-scale CBS systems maintain trillions of block-mapping key-value entries as metadata to track the storage location of each virtual block. Although SSD-based KV stores have been widely adopted in cloud systems for their high efficiency and durability, current SSD-based schemes face significant challenges in achieving deterministic access latency for latency-sensitive metadata services. Our experimental observations indicate that the substantial long-tail latency is primarily caused by (1) I/O blocking due to internal tasks of SSDs including modern Zone Namespace SSDs; and (2) additional disk I/Os when querying high-level indexes across memory and SSDs under memory-constrained environments. In this article, we propose an SSD-based SIndex to store trillions of block-mapping entries for latency-critical cloud block storage, which performs comprehensive latency optimization across storage I/O scheduling and high-level indexing. To prevent long-tail I/Os while avoiding intrusive device modifications, SIndex introduces an inter-SSD I/O scheduling mechanism based on read/write separation and SSD state transitions, which mitigates latency fluctuations induced by garbage collection on conventional SSDs and zone operations on Zone Namespace SSDs. Additionally, SIndex employs opportunistic I/O speculation and a concurrent request balancing mechanism to reduce read disturbance and I/O contention. To query the storage location of targeted block-mapping entries with bounded latency, SIndex proposes a memory-efficient high-level index incorporates with a static data layout, preventing time-consuming disk lookups by keeping the index in memory. We evaluate the SIndex prototype using a variety of benchmarks and real-world traces on commodity SSDs. The results demonstrate that SIndex outperforms RocksDB and other approaches by up to 11.4× in tail latency, keeping the 99.99th-percentile latency below 400 μs. Shucheng Wang, Zhandong Guo, Kaiye Zhou, Jun Xu 0037, Qiang Cao 0001 |
ACM Trans. Storage | 1 |
| 2025 | LCache: Log-Structured SSD Caching for Training Deep Learning ModelsabstractTraining deep learning models is computationally demanding and data-intensive. Existing approaches utilize local SSDs within training servers to cache datasets, thereby accelerating data loading during model training. However, we experimentally observe that data loading remains a performance bottleneck when randomly retrieving small-sized sample files on SSDs. In this paper, we introduce LCache, a log-structured dataset caching mechanism designed to fully leverage the I/O capabilities of SSDs and reduce I/O-induced training stalls. LCache determines the randomized dataset access order by extracting the pseudo-random seed from the training frameworks. It then aggregates small-sized sample files into larger chunks and stores them in a log file on SSDs, thus enabling sequential I/O requests on data retrieval and improving data loading throughput. Further-more, LCache proposes a real-time log reordering mechanism that strategically schedules cached data to organize logs across different epochs, which enhances cache utilization and minimizes data retrieval from low-performance remote storage systems. Additionally, LCache incorporates an MetaIndex to enable rapid log traversal and querying. We evaluate LCache with various real-world DL models and datasets. LCache outperforms the native PyTorch Dataloader and NoPFS by up to 9.4x and 7.8x in throughput,. respectively, Shucheng Wang, Zhandong Guo, Jian Sheng, Kaiye Zhou, Qiang Cao 0001 |
DATE | 1 |
| 2025 | Comprehensive performance evaluation of valuable medical equipment based on cloud modelling and combined weighting methodologies
Xingtong Zhang, Saifeng Fang, Yongchun Jin, Shucheng Wang, Yunhua Xu |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | ParaCkpt: Heterogeneous Multi-Path Checkpointing Mechanism for Training Deep Learning ModelsabstractTraining large deep learning models is extremely computationally intensive and time-consuming; therefore, it relies on checkpointing mechanisms to save snapshots promptly, ensuring rapid recovery from a myriad of failures. Existing checkpointing approaches save snapshots to either CPU memory or storage, overlooking their aggregated I/O capability. In this paper, we propose a heterogeneous multi-path checkpointing mechanism, ParaCkpt, to make full use of both PCle-bandwidth and I/O capability of memory and storage to accelerate check-pointing. ParaCkpt first identifies multiple paths for GPUs to CPU memory, local and remote storages, and determines their available bandwidths. Then, ParaCkpt strategically partitions the training model states into a set of path-based shards and drains them from GPUs to the memory and storage in parallel. Moreover, ParaCkpt employs a two-stage persistence strategy to flush in-memory shards to local SSDs in the background, and then stores local shards using compression to remote storage. Finally, ParaCkpt maintains global snapshots distributed across memory and storage, enabling rapid recovery via the multi-path way. We evaluate ParaCkpt with various real-world deep learning models. ParaCkpt outperforms native Pytorch and state-of-the-art asynchronous checkpointing approaches by up to 96 x and 2.3 x in throughput, respectively. Shucheng Wang, Qiang Cao 0001, Kaiye Zhou, Jun Xu 0037, Zhandong Guo, Jiannan Guo 0001 |
ICCD | 1 |
| 2024 | SIndex: An SSD-based Large-scale Indexing with Deterministic Latency for Cloud Block StorageabstractThe Solid State Drives (SSD) based key-value stores face significant challenges in achieving deterministic access latency. We experimentally observe the long-tail latency is mainly caused by I/O blocking induced by SSD’s internal tasks. In this paper, we propose an SSD-based SIndex to store hundreds of billions of block-mapping entries for the latency-critical cloud block storage. To hide the latency fluctuations induced by garbage collection and buffer flushing, SIndex proposes an inter-SSD I/O scheduling based on read/write separation and SSD state transition, while adopting opportunistic request speculation and balancing mechanism to mitigate read disturbance and I/O contention. Moreover, SIndex introduces a write-staging buffer cache and a two-stage sync mechanism to preferentially buffer updated data before synchronizing them to SSDs. We evaluate the SIndex prototype with a variety of benchmarks and real-world traces on commodity SSDs. SIndex is demonstrated to outperform RocksDB and other approaches by up to 11.2 × in tail latency without affecting the throughput performance. Shucheng Wang, Kaiye Zhou, Zhandong Guo, Qiang Cao 0001, Jun Xu 0037, Jie Yao 0001 |
ICPP | 1 |
| 2024 | Explorations and Exploitation for Parity-based RAIDs with Ultra-fast SSDsabstractFollowing a conventional design principle that pays more fast-CPU-cycles for fewer slow-I/Os, popular software storage architecture Linux Multiple-Disk (MD) for parity-based RAID (e.g., RAID5 and RAID6) assigns one or more centralized worker threads to efficiently process all user requests based on multi-stage asynchronous control and global data structures, successfully exploiting characteristics of slow devices, e.g., Hard Disk Drives (HDDs). However, we observe that, with high-performance NVMe-based Solid State Drives (SSDs), even the recently added multi-worker processing mode in MD achieves only limited performance gain because of the severe lock contentions under intensive write workloads. In this paper, we propose a novel stripe-threaded RAID architecture, StRAID, assigning a dedicated worker thread for each stripe-write (one-for-one model) to sufficiently exploit high parallelism inherent among RAID stripes, multi-core processors, and SSDs. For the notoriously performance-punishing partial-stripe writes that induce extra read and write I/Os, StRAID presents a two-stage stripe write mechanism and a two-dimensional multi-log SSD buffer. All writes first are opportunistically batched in memory, and then are written into the primary RAID for aggregated full-stripe writes or conditionally redirected to the buffer for partial-stripe writes. These buffered data are strategically reclaimed to the primary RAID. We evaluate a StRAID prototype with a variety of benchmarks and real-world traces. StRAID is demonstrated to outperform MD by up to 5.8 times in write throughput. Shucheng Wang, Qiang Cao 0001, Hong Jiang 0001, Ziyi Lu, Jie Yao 0001, Yuxing Chen 0003, Anqun Pan |
ACM Trans. Storage | 1 |
| 2023 | PMLDS: An LSM-Tree Direct Managed Storage for Key-Value Stores on Byte-Addressable DevicesabstractExisting key-value stores (KVSs) based on log-structured merge-tree (LSM-tree) have been broadly deployed in practice to leverage characteristics of conventional block storage via file system, but lack effective exploitation for emerging byte-addressed persistent memory (PM). We reveal that these KVSs running upon existing PM-aware File systems cause inefficient PM I/O behaviors, including 1) numerous page faults, 2) I/O misaligned with cacheline, and 3) bandwidth wastage of concurrent I/O threads. To make full use of PM without major modification for existing LSM-based KVSs, this paper proposes PMLDS, a direct managed storage for LSM-tree-based KVSs directly running upon PM. PMLDS acts as a unified I/O layer to handle all requests from KVS to PM. PMLDS designs an LSM-tree-aware data layout to directly map the KVS’s persistent objects to the storage slots with fixed location and size, thus simplifying and replacing the file system’s functionality with a minor modification. To improve I/O efficiency, PMLDS further presents three key techniques: 1) pre-allocating reusable data slots to avoid page faults, 2) forcing cacheline-alignment for small requests, and 3) scheduling asynchronous I/O threads to harness PM’s limited parallelism. We implement PMLDS and evaluate it with popular RocksDB under a variety of workloads. The results show that compared to representative PM-aware file systems such as Ext4-DAX, XFS-DAX, NOVA, and WineFS, PMLDS improves the write performance of RocksDB by up to 2.1 × while reducing the read latency by 20%~50%. Ziyi Lu, Qiang Cao 0001, Shucheng Wang, Jie Yao 0001, Xiangrui Yang 0001 |
ICPP | 3 |
| 2022 | PATS: Taming Bandwidth Contention between Persistent and Dynamic MemoriesabstractEmerging persistent memory (PM) with fast per-sistence and byte-addressability physically shares the memory channel with DRAM-based main memory. We experimentally uncover that the throughput of application accessing DRAM collapses when multiple threads access PM due to head-of-line blockage in the memory controller within CPU. To address this problem, we design a PM-Accessing Thread Scheduling (PATS) mechanism that is guided by a contention model, to adaptively tune the maximum number of contention-free concurrent PM-threads. Experimental results show that even with 14 concurrent threads accessing PM, PATS is able to allow only up to 8% decrease in the DRAM-throughput of the front-end applications (e.g., Memcached), gaining 1.5x PM-throughput speedup over the default configuration. Shucheng Wang, Qiang Cao 0001, Ziyi Lu, Hong Jiang 0001 |
DATE | 1 |
| 2022 | p2KVS: a portable 2-dimensional parallelizing framework to improve scalability of key-value stores on SSDsabstractAttempts to improve the performance of key-value stores (KVS) by replacing the slow Hard Disk Drives (HDDs) with much faster Solid-State Drives (SSDs) have consistently fallen short of the performance gains implied by the large speed gap between SSDs and HDDs, especially for small KV items. We experimentally and holistically explore the root causes of performance inefficiency of existing LSM-tree based KVSs running on powerful modern hardware with multicore processors and fast SSDs. Our findings reveal that the global write-ahead-logging (WAL) and index-updating (MemTable) can become bottlenecks that are as fundamental and severe as the commonly known LSM-tree compaction bottleneck, under both the single-threaded and multi-threaded execution environments. Ziyi Lu, Qiang Cao 0001, Hong Jiang 0001, Shucheng Wang |
EuroSys | 4 |
| 2022 | Mlog: Multi-log Write Buffer upon Ultra-fast SSD RAIDabstractParity-based RAID suffering from partial-stripe write-penalty has to introduce write buffer to fast absorb and merge incoming writes, and then flush them to RAID array in batch. However, we experimentally observe that the popular buffering mechanism as Linux RAID journal and partial parity logging (PPL) becomes a bottleneck for ultra-fast SSD-based RAID, and we further uncover that the centralized log-buffer model is the prime cause. Shucheng Wang, Qiang Cao 0001, Ziyi Lu, Jie Yao 0001 |
ICPP | 1 |
| 2022 | StRAID: Stripe-threaded Architecture for Parity-based RAIDs with Ultra-fast SSDs
Shucheng Wang, Qiang Cao 0001, Ziyi Lu, Hong Jiang 0001, Jie Yao 0001 |
USENIX ATC | 1 |
| 2022 | Exploration and Exploitation for Buffer-Controlled HDD-Writes for SSD-HDD Hybrid Storage ServerabstractHybrid storage servers combining solid-state drives (SSDs) and hard-drive disks (HDDs) provide cost-effectiveness and μs-level responsiveness for applications. However, observations from cloud storage system Pangu manifest that HDDs are often underutilized while SSDs are overused, especially under intensive writes. It leads to fast wear-out and high tail latency to SSDs. On the other hand, our experimental study reveals that a series of sequential and continuous writes to HDDs exhibit a periodic, staircase-shaped pattern of write latency, i.e., low (e.g., 35 μs), middle (e.g., 55 μs), and high latency (e.g., 12 ms), resulting from buffered writes within HDD’s controller. It inspires us to explore and exploit the potential μs-level IO delay of HDDs to absorb excessive SSD writes without performance degradation. We first build an HDD writing model for describing the staircase behavior and design a profiling process to initialize and dynamically recalibrate the model parameters. Then, we propose a Buffer-Controlled Write approach (BCW) to proactively control buffered writes so that low- and mid-latency periods are scheduled with application data and high-latency periods are filled with padded data. Leveraging BCW, we design a mixed IO scheduler (MIOS) to adaptively steer incoming data to SSDs and HDDs. A multi-HDD scheduling is further designed to minimize HDD-write latency. We perform extensive evaluations under production workloads and benchmarks. The results show that MIOS removes up to 93% amount of data written to SSDs, reduces average and 99 th -percentile latencies of the hybrid server by 65% and 85%, respectively. Shucheng Wang, Ziyi Lu, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang, Changsheng Xie 0001 |
ACM Trans. Storage | 1 |
| 2021 | EFLOG: A Full Stream-Logging Scheme with Erasure Coding in Cloud Storage SystemsabstractLarge-scale cloud storage systems use the logging mechanism to sequentially write data in an append-only manner. The write stream needs to be first appended and persisted into logging files, and then encoded with erasure coding (EC) in underlying storage. This introduces significant overhead to small write operations. To solve this problem, we propose EFLOG, a full-streaming storage framework that combines Logging and inter-log EC mechanisms. EFLOG evenly schedules front-end write streams across log files in each disk with append-only manner. In background, EFLOG determines unprotected logged data and seals them into ECblocks. Afterwards, EFLOG concurrently encodes data with multi-threads and stores parity data into parity disks. Results of our trace-driven evaluation show that, EFLOG can achieve up to 1.01GB/s write throughput with RS(4, 2) codes built upon 6 SSD disks. Qiang Cao 0001, Shucheng Wang, Changsheng Xie 0001 |
NAS | 3 |
| 2020 | BCW: Buffer-Controlled Writes to HDDs for SSD-HDD Hybrid Storage Server
Shucheng Wang, Ziyi Lu, Qiang Cao 0001, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang |
FAST | 1 |
| 2020 | SeRW: Adaptively Separating Read and Write upon SSDs of Hybrid Storage Server in CloudsabstractNowadays, cloud providers embrace hybrid storage servers to reap both high IO performance of solid-state drives (SSDs) and low-cost of hard disk drives (HDDs). These hybrid storage servers generally employ SSDs as primary storage directly serving requests from front-end applications while using HDDs as the secondary storage to provide sufficient storage capacity. Qiang Cao 0001, Shucheng Wang, Jie Yao 0001, Puyuan Yang |
ICPP | 3 |
| 2019 | Analysis of and Optimization for Write-dominated Hybrid Storage Nodes in CloudabstractCloud providers like the Alibaba cloud routinely and widely employ hybrid storage nodes composed of solid-state drives (SSDs) and hard disk drives (HDDs), reaping their respective benefits: performance from SSD and capacity from HDD. These hybrid storage nodes generally write incoming data to its SSDs and then flush them to their HDD counterparts, referred to as the SSD Write Back (SWB) mode, thereby ensuring low write latency. When comprehensively analyzing real production workloads from Pangu, a large-scale storage platform underlying the Alibaba cloud, we find that (1) there exist many write dominated storage nodes (WSNs); however, (2) under the SWB mode, the SSDs of these WSNs suffer from severely high write intensity and long tail latency. To address these unique observed problems of WSNs, we present SSD Write Redirect (SWR), a runtime IO scheduling mechanism for WSNs. SWR judiciously and selectively forwards some or all SSD-writes to HDDs, adapting to runtime conditions. By effectively offloading the right amount of write IOs from overburdened SSDs to underutilized HDDs in WSNs, SWR is able to adequately alleviate the aforementioned problems suffered by WSNs. This significantly improves overall system performance and SSD endurance. Our trace-driven evaluation of SWR, through replaying production workload traces collected from the Alibaba cloud in our cloud testbed, shows that SWR decreases the average and 99til-percentile latencies of SSD-writes by up to 13% and 47% respectively, notably improving system performance. Meanwhile the amount of data written to SSDs is reduced by up to 70%, significantly improving SSD lifetime. Shucheng Wang, Qiang Cao 0001, Ziyi Lu, Hong Jiang 0001, Jie Yao 0001, Puyuan Yang |
SoCC | 2 |
| 2018 | Density-Based Fuzzy C-Means Multi-center Re-clustering Radar Signal Sorting AlgorithmabstractAs the improving strategic position of electronic warfare in modern warfare, radar sorting detection becomes the eye of modern information warfare and plays an important role in it. This paper designs a new pulse radar sorting algorithm: a Density-Based Fuzzy C-Means Multi-Center Re-Clustering (DFCMRC) radar signal sorting algorithm. This algorithm mainly combines the advantages of Density-Based Spatial Clustering of Applications with Noise (DBSCAN) clustering algorithm and Fuzzy C-means (FCM) clustering algorithm. This paper also optimizes the structure of the DFCMRC algorithm, which changes the algorithm that randomly generated the initial center point to the Clustering by Fast Search and Find of Density Peaks (CFSFDP) algorithm. After comparison tests, the DFCMRC algorithm sorting result is better than the K-means algorithm, the DBSCAN algorithm and the FCM algorithm. Also, the membership grade description of DFCMRC makes more sense than the FCM's. Accelerated optimized DFCMRC algorithm can reduce more than half iterations, which greatly shortens the algorithm calculation time. Shucheng Wang |
ICMLA | 2 |
| 2018 | Design of River Water Quality Assessment and Prediction AlgorithmabstractDue to the rapid population growth and economic development, water environmental protection pressures has been increasing recently. This paper focuses on the pollution of water quality, building a water quality assessment model to analyze the water quality level, and makes an objective further prediction of the trend of its factors. In this paper, the mutation factor of genetic algorithm is introduced into the PSO algorithm. The Least Squares Support Vector Machine (LS-SVM) based on adaptive Particle Swarm Optimization (PSO) algorithm used to optimize the hyper-parameter builds one water quality classification assessment model. The fuzzy information granulation method is combined with the Least Square Support Regression (LS-SVR) to set up a water quality time series model, which can predict the trend of changes in water quality data in three days. With the help of the theoretical analysis and experimental data, this assessment model and the prediction algorithm are faster in training speed and higher in accuracy, compared with the traditional BP neural network. Shucheng Wang |
ICMLA | 2 |