Shuhan Bai

dblp:327/1851 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-0842-7602ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-author · 10 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hitchhike: Efficient Request Submission via Deferred Enforcement of Address Contiguity
abstract
Modern storage systems operate under high concurrency, making large volumes of outstanding I/Os the norm. However, current I/O submission logic requires requests assigned to the same CPU core to be processed in a serialized manner, turning the software stack into a bottleneck due to high per-request overhead.
Xuda Zheng, Jian Zhou 0004, Shuhan Bai, Runjin Wu, Xianlin Tang, Hong Jiang 0001, Fei Wu 0005
ASPLOS (2)3
2026 REMUS: Efficient Multirequest Scheduling in Computational Storage Devices
abstract
Numerous data-intensive applications benefit from offloading data processing to storage devices, typically Computational Storage Devices (CSD). Co-locating requests from diverse applications to one CSD offers better performance and power efficiency than dedicating CSDs to a single application. However, current CSD scheduling frameworks struggle to effectively manage contention among CPU, flash I/O, and buffer resources across requests due to less consideration in their inter-dependencies. This paper proposes REMUS, a CSD scheduling framework handling multiple requests for commercial SSD with multiple homogeneous cores. The key idea of REMUS is to allocate workloads across multiple cores based on the distribution of the Logical Block Address (LBA) of requests, and to mitigate stall time by sorting requests according to their urgency for resources, where urgency is quantified by each request’s remaining buffer capacity. Furthermore, a request batching scheme that intelligently groups the requests to be scheduled according to their characteristics is proposed to provide congestion control for REMUS and to minimize the contention it introduces. We conduct experiments on both a simulator and a real CSD platform. The experiment results show that REMUS improved throughput by 1.51× on the simulator and 1.39× on the real platform on average compared to the baselines.
Yun Huang 0005, Shuhan Bai, Heng-Lin Yen, Nan Guan, Tei-Wei Kuo, Chun Jason Xue
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 Fraggle: Reducing File Fragmentation on DRAM-Less SSD
abstract
Mobile devices adopt DRAM-less SSDs as the flash storage to reduce power consumption and manufacturing costs, yet DRAM-less SSDs are significantly impacted by file fragmentation, thereby degrading the responsiveness of mobile devices. Existing methods to mitigate file fragmentation, including preallocation and defragmentation, face critical limitations. Preallocation prevents file fragmentation by reserving contiguous free LBA space for individual files. However, within the limited LBA space exposed by the SSD, contiguous free space is gradually consumed, and fragmented by scattered data blocks, reducing the effectiveness of preallocation. Defragmentation recovers file access performance by reorganizing fragmented file layouts into contiguous ones. However, when applied to DRAM-less SSD, defragmentation introduces excessive cache eviction of dirty mapping entries to the slow NAND flash, causing significant performance disruption. This paper presents Fraggle, a host-device co-design that mitigates file fragmentation on DRAM-less SSDs by preserving the efficacy of preallocation within an extended LBA space. To effectively manage the extended LBA space, Fraggle introduces two key components: (1) HashFTL, a compact, cache-efficient FTL mapping mechanism specifically optimized for DRAM-less SSDs, and (2) a device-assisted space allocator that simultaneously reduces space allocation overhead and resolves layout imbalance issues in HashFTL caused by sparse LBA distribution. Evaluation results demonstrate that, compared to F2FS and state-of-the-art preallocation methods, Fraggle accelerates execution time by up to 9.26× and 8.07× under SQLite and real-world mobile application workloads.
Weizhou Huang, Hepei Wu, Shuhan Bai, Jian Zhou 0004, Fei Wu 0005
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 ReCoW: Kernel-User Collaborative Copy-on-Write Transactions for Persistent Memory via Address Remapping
abstract
Copy-on-write (CoW) is a widely used technique to enable failure-atomic transactions for persistent memory (PM), which avoids the double-write problem in logging-based transactions. However, balancing copy granularity and metadata size for large objects is challenging. Oversized metadata introduces nontrivial overhead in tracking the out-of-place updated data blocks. To the best of our knowledge, prior work has not adequately addressed this metadata issue. This paper presents ReCoW, a kernel-user collaborative CoW transaction system that employs the hardware memory management unit (MMU) to accelerate the data referencing. ReCoW leverages virtual memory remapping to consolidate fragmented data into a contiguous virtual address space, thus significantly reducing the metadata sizes, which in turn substantially decreases metadata performance overhead. Our evaluation under real-world workloads indicates that ReCoW outperforms state-of-the-art competitors, PMDK, SpecPMT, and ArchTM by 3.90x, 2.76x, and 1.45x on average in throughput, respectively, while reducing metadata size to about 1% of that of comparable systems.
Aoxin Wei, Jian Zhou 0004, Shuhan Bai, Yufan Jia, Jintian Wu, Fei Wu 0005, Hong Jiang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2025 Locality-Aware Data Placement for NUMA Architectures: Data Decoupling and Asynchronous Replication
abstract
Non-Uniform Memory Access (NUMA) architectures bring new opportunities and challenges to bridge the gap between computing power and memory performance. Their complex memory hierarchies feature non-uniform access performance, known as NUMA locality, indicating data placement and access without NUMA-awareness significantly impact performance. Existing NUMA-aware solutions often prioritize fast local access but at the cost of heavy replication overhead, suffering a read-write performance tradeoff and limited scalability. To overcome these limitations, this paper presents Ladapa, a scalable and high-performance locality-aware data placement strategy. The key insight is decoupling data into metadata and data layers, allowing independent management with adaptive asynchronous replication for lower overhead. Additionally, Ladapa employs multi-level metadata management leveraging fast caches for efficient data location, further boosting performance. Experimental results show that Ladapa outperforms typical replication techniques by up to 27.37× in write performance and 1.63× in read performance.
Shuhan Bai, Haowen Luo, Burong Dong, Jian Zhou 0004, Fei Wu 0005
DATE1
2025 REMUS: Efficient Multi-Request Scheduling in Computational Storage Devices
Yun Huang 0005, Shuhan Bai, Heng-Lin Yen, Nan Guan, Tei-Wei Kuo, Xue (Steve) Liu, Chun Jason Xue
RTCSA2
2025 AquaPipe: A Quality-Aware Pipeline for Knowledge Retrieval and Large Language Models
abstract
The knowledge retrieval methods such as Approximate Nearest Neighbor Search (ANNS) significantly enhance the generation quality of Large Language Models (LLMs) by introducing external knowledge, and this method is called Retrieval-augmented generation (RAG). However, due to the rapid growth of data size, ANNS tends to store large-scale data on disk, which greatly increases the response time of RAG systems. This paper presents AquaPipe, which pipelines the execution of disk-based ANNS and the LLM prefill phase in an RAG system, effectively overlapping the latency of knowledge retrieval and model inference to enhance the overall performance, while guaranteeing data quality. First, ANNS's recall-aware prefetching strategy enables the early return of partial text with acceptable accuracy so the prefill phase can launch before getting the full results. Then, we adaptively choose the remove-after-prefill or re-prefill strategies based on the LLM cost model to effectively correct disturbed pipelines caused by wrong early returns. Finally, the pipelined prefill dynamically changes the granularity of chunk size to balance the overlap efficiency and GPU efficiency, adjusting to ANNS tasks that converge at different speeds. Our experiments have demonstrated the effectiveness of AquaPipe. It successfully masks the latency of disk-based ANNS by 56% to 99%, resulting in a 1.3× to 2.6× reduction of the response time of the RAG, while the extra recall loss caused by prefetching is limited to approximately 1%.
Runjie Yu, Weizhou Huang, Shuhan Bai, Jian Zhou 0004, Fei Wu 0005
Proc. ACM Manag. Data3
2025 Lemonade: Learning-based Heterogeneous Metadata Offloading for Disaggregated Memory
abstract
Direct Access (DA) in Disaggregated Memory (DM) is a promising solution that meets the high-performance requirements of AI applications. However, it lacks effective support for metadata management, making metadata operations the major bottleneck. To address this, we propose Lemonade, a l earning-based h e terogeneous m etadata o ffloadi n g for dis a ggregate d m e mory. Lemonade splits the metadata into highly regular and irregular ones, thus offloading the former into the client to avoid remote queries and enabling request redirection in the SmartNIC for the latter to ensure cost-effective correction and updates. Evaluations under microbenchmark and YCSB workloads indicate that Lemonade reduces latency by 72.8% and achieves a 1.43× increase in throughput compared to the state-of-the-art systems.
Zeming Ma, Jian Zhou 0004, Xiaochang Ma, Shuhan Bai, Fei Wu 0005
ACM Trans. Embed. Comput. Syst.5
2024 LaVA: An Effective Layer Variation Aware Bad Block Management for 3D CT NAND Flash
abstract
3D NAND flash with charge trap (CT) technology has been developed by stacking multiple layers vertically to boost storage capacity while ensuring reliability and scalability. One of its critical characteristics is the large endurance variation among and inside blocks and layers. With this feature, traditional bad block management (BBM), which determines block lifetime by the page with worst endurance, results in underutilization of solid state drive (SSD) usage. In this paper, a layer variation aware and fault-tolerant bad block management, named LaVA, is proposed to prolong the lifetime of 3D NAND flash storage. The relevant layer, instead of the entire flash block, is discarded at a finer granularity when a page failure is encountered. Experimental results based on real-world workloads show that LaVA can significantly extend the endurance of 3D CT NAND flash (30.6%-62.2%) with a small performance degradation (less than 10% increase of tail I/Oresponse time), compared to the conventional technique.
Shuhan Bai, You Zhou 0009, Fei Wu 0005, Changsheng Xie 0001, Tei-Wei Kuo, Chun Jason Xue
DATE1
2023 SERICO: Scheduling Real-Time I/O Requests in Computational Storage Drives
abstract
The latency and energy consumption incurred by I/O accesses are significant in data-centric computing systems. Computational Storage Drive (CSD) can largely reduce data movement, and thus reduce I/O latency and energy consumption by offloading data-intensive processing to processors inside the storage device. In this paper, we study the problem of how to efficiently utilize the limited processing and memory resources of CSD to simultaneously serve multiple I/O requests from various applications with different real-time requirements. We proposed SERICO, a system of scheduling computational I/O requests in CSD. The key idea of SERICO is to perform admission control of real-time computational I/O requests by online schedulability analysis, to avoid wasting the processing resources and memory capacity of CSD in doing meaningless work for those requests deemed to violate the timing constraints. Each admitted computational I/O request is served in a controlled manner with carefully designed parameters, to meet its timing constraint with minimal memory cost. We evaluate SERICO with both synthetic workloads on simulators and representative applications on realistic CSD hardware. Experiment results show that SERICO significantly outperforms the default method used in the CSD device and the standard deadline-driven scheduling approach.
Yun Huang 0005, Nan Guan, Shuhan Bai, Tei-Wei Kuo, Chun Jason Xue
DATE3
2023 Pipette: Efficient Fine-Grained Reads for SSDs
abstract
Big data applications, such as recommendation system and social network, often generate a huge number of fine-grained reads to the storage. Block-oriented storage devices upon the traditional storage system rely on the paging mechanism to migrate pages to the host DRAM, tending to suffer from these fine-grained read operations in terms of I/O traffic as well as performance. Motivated by this challenge, an efficient fine-grained read framework, Pipette, is proposed in this article as an extension to the traditional I/O framework. With adaptive design for caching, merging, and scheduling, Pipette explores locality and acceleration for fine-grained read requests to establish an efficient byte-granular read path upon the dedicated byte-addressable interface. When the Pipette prototype on an SSD runs popular workloads, we measured throughput gains by up to 50% and 54% with traffic reduction in the range of$41.3\times $and$56.5\times $.
Shuhan Bai, Hu Wan 0001, Yun Huang 0005, Xuan Sun 0003, Fei Wu 0005, Changsheng Xie 0001, Hung-Chih Hsieh, Tei-Wei Kuo, Chun Jason Xue
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Pipette: efficient fine-grained reads for SSDs
abstract
Big data applications, such as recommendation system and social network, often generate a huge number of fine-grained reads to the storage. Block-oriented storage devices tend to suffer from these fine-grained read operations in terms of I/O traffic as well as performance. Motivated by this challenge, a fine-grained read framework, Pipette, is proposed in this paper, as an extension to the traditional I/O framework. With an adaptive caching design, Pipette framework offers a tremendous reduction in I/O traffic as well as achieves significant performance gain. A Pipette prototype was implemented with Ext4 file system on an SSD for two real-world applications, where the I/O throughput is improved by 31.6% and 33.5%, and the I/O traffic is reduced by 95.6% and 93.6%, respectively.
Shuhan Bai, Hu Wan 0001, Yun Huang 0005, Xuan Sun 0003, Fei Wu 0005, Changsheng Xie 0001, Hung-Chih Hsieh, Tei-Wei Kuo, Chun Jason Xue
DAC1