Shun Gai

dblp:329/0949 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-7967-0036ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 zBuffer: Zero-Copy and Metadata-Free Serialization for Fast RPC with Scatter-Gather Reflection
abstract
This paper presents zBuffer, a zero-copy and metadata-free serialization library for high-performance and low-cost RPCs. At the core of zBuffer is scatter-gather reflection, a novel technique that collaboratively (i) leverages the NIC scatter-gather hardware feature to offload the costly data coalescing, and (ii) utilizes the static reflection mechanism of modern programming languages to enable type queries on complex data objects without requiring explicit metadata construction. We leverage C++ language features, mainly including template meta-programming and macros, to realize static reflection at compile time. Based on zBuffer, we design a fast RPC system (called zRPC) which eliminates all RPC memory copy overheads not only in (de)serialization but also in network transmission. Extensive evaluation shows that zBuffer/zRPC significantly outperforms state-of-the-art serialization/RPC mechanisms: zBuffer is approximately 7× faster than Cornflakes in serialization for complex data objects; and zRPC reduces 99th percentile latency by 21% and achieves 62% higher throughput than eRPC on the Masstree key-value (KV) store with the YCSB benchmark.
Huiba Li, Shun Gai, Youmin Chen, Yiming Zhang 0003
PPoPP3
2025 Cheetah: Metadata Aggregation for Fast Object Storage without Distributed Ordering
abstract
Object stores usually maintain the mapping of objects to data servers' disk volumes (referred to as volume metadata) in a central directory, while storing the object data's in-volume offsets (referred to as offset metadata) together with the data on data servers. Unfortunately, the separation between volume/offset metadata complicates the processing of an object put: to ensure consistency, the multiple writes of the object's volume/offset metadata and object data have to be orchestrated in a particular order, which severely lowers object I/O performance. We propose a write-optimal structure called MetaX that aggregates all metadata of a put, including both volume and offset metadata as well as other meta information such as data checksum and temporary meta-log. Based on MetaX, we design the Cheetah object store, which organizes object storage into rich metadata storage (on meta servers) and raw data storage (on data servers). Cheetah removes the distributed ordering constraint on the multiple metadata/data writes by enforcing local atomicity of writing MetaX, while still ensuring consistency. Evaluation shows that Cheetah significantly outperforms existing object stores.
Yiming Zhang 0003, Li Wang 0123, Shengyun Liu, Shun Gai, Xin Yao 0008, Kai Chen 0005, Dongsheng Li 0001, Jiwu Shu
EuroSys4
2023 HSS: A Hierarchical Semantic Similarity Hard Negative Sampling Method for Dense Retrievers
Xinjia Xie, Shun Gai, Zhen Huang 0006, Minghao Hu 0001, Ankun Wang
MMM (2)3
2023 SRockDB: A Range-Query Optimized Database Based on RocksDB
abstract
Data stores based on Log-Structured Merge Tree (LSM-Tree) are widely used in data centres, Artificial Learning and Machine Learning. As the core data structure, LSM-Tree is efficient in writing operations due to the out-of-place update and leveled design. However, it has limitations in writing, reading and space amplifications which decrease performance, especially range-query. Range query is one of the most important operations in data processing. To address this problem, we propose SRockDB. SRockDB is specifically designed to improve the performance of range queries by the in-memory cache. Compared to the block cache in RocksDB, SRockDB is more effective and coalesces adjacent keys to cache more data in limited memory. In experiments, SRockDB achieves 2.76× Queries Per Second (QPS), up to 41.74% lower tail latency, without downsides on the performance of write and point read.
Shun Gai, Xinjia Xie
SMC1
2023 UrsaX: Integrating Block I/O and Message Transfer for Ultrafast Block Storage on Supercomputers
abstract
It is increasingly important for the next-generation exascale supercomputers to extend its applications beyond traditional high-performance computing (HPC) scenarios, so as to achieve high social and economic benefit. Similar to Amazon Web Services (AWS) and Alibaba Cloud, cloud-style virtual HPC service is a promising application scenario on supercomputers, for which remote block storage is the key to provide tenants with supercomputers’ extremely high storage performance. Unfortunately, the state-of-the-art block storage software systems (such as URSA and Ceph) cannot adapt to the advanced hardware features of supercomputers. This article presents UrsaX, an efficient block storage service for our next-generation Tianhe exascale supercomputer that is equipped with the high-performance global express (GLEX) network and nonvolatile memory express (NVMe) SSDs. UrsaX’s virtual disks, which can be mounted like normal physical ones, enable not only traditional HPC applications but also supercomputer-oblivious POSIX applications to enjoy the high performance of supercomputers. At the core of UrsaX is with a novel design of the efficient integration of on-disk block I/O and in-network message transfer on supercomputers. UrsaX utilizes the NVMe Fabrics kernel module to expand the NVMe standard on the supercomputer network, and separates metadata I/O and data I/O of blocks, respectively, being handled over the mini packet (MP) and remote direct memory access (RDMA) protocols. We thoroughly explore the design space for remote block storage on supercomputers, including parallelism, scalability, fault tolerance, and consistency. We conduct an extensive evaluation on a subset of our exascale supercomputer consisting of 44 storage machines (each with four NVMe SSDs). The result shows that UrsaX achieves local-storage-level I/O latency (tens of microseconds) while being able to linearly increase the aggregate performance (IOPS and throughput) as the system scale increases, an order of magnitude higher than the state-of-the-art block storage systems.
Shun Gai, Yiming Zhang 0003, Xuchao Xie, Yong Dong, Zhenlong Song
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Oasis: Controlling Data Migration in Expansion of Object-based Storage Systems
abstract
Object-based storage systems have been widely used for various scenarios such as file storage, block storage, blob (e.g., large videos) storage, and so on, where the data is placed among a large number of object storage devices (OSDs). Data placement is critical for the scalability of decentralized object-based storage systems. The state-of-the-art CRUSH placement method is a decentralized algorithm that deterministically places object replicas onto storage devices without relying on a central directory. While enjoying the benefits of decentralization such as high scalability, robustness, and performance, CRUSH-based storage systems suffer from uncontrolled data migration when expanding the capacity of the storage clusters (i.e., adding new OSDs), which is determined by the nature of CRUSH and will cause significant performance degradation when the expansion is nontrivial. This article presents MapX , a novel extension to CRUSH that uses an extra time-dimension mapping (from object creation times to cluster expansion times) for controlling data migration after cluster expansions. Each expansion is viewed as a new layer of the CRUSH map represented by a virtual node beneath the CRUSH root. MapX controls the mapping from objects onto layers by manipulating the timestamps of the intermediate placement groups (PGs). MapX is applicable to a large variety of object-based storage scenarios where object timestamps can be maintained as higher-level metadata. We have applied MapX to the state-of-the-art Ceph-RBD (RADOS Block Device) to implement a migration-controllable, decentralized object-based block store (called Oasis ). Oasis extends the RBD metadata structure to maintain and retrieve approximate object creation times (for migration control) at the granularity of expansion layers. Experimental results show that the MapX -based Oasis block store outperforms the CRUSH-based Ceph-RBD (which is busy in migrating objects after expansions) by 3.17× ∼ 4.31× in tail latency, and 76.3% (respectively, 83.8%) in IOPS for reads (respectively, writes).
Yiming Zhang 0003, Li Wang 0152, Shun Gai, Qiwen Ke, Zhenlong Song, Guangtao Xue, Jiwu Shu
ACM Trans. Storage3
2022 LSTMcon: A Novel System of Portfolio Management Based on Feedback LSTM with Confidence
abstract
Trading carries a substantial amount of risk and making adequately informed decisions cannot be overemphasized.In order to propose a more reasonable strategy on portfolio arrangement, we design LSTMcon, a two-stage system that consists of a assets price prediction model and a decisionmaking strategy based on ensemble rules.As for next-day price prediction, we implement an LSTM model with feedback mechanism and devise a series of training settings.The feedback mechanism uses the deviation between predicted price and actual price to correct the prediction result from LSTM.To decrease the transaction cost, we design a three-day trading period and adopt an iterative prediction approach.Our model achieves the accuracy of 98.5% on GOLD and 98.8% on BTC finally.In addition, we devise a decision-making system after getting the predicted data.We modify the predicted price by giving everyone a certain confidence level based on three approaches (reward and punishment mechanism, sequential days rules, historical price relying).We combine these rules and give a comprehensive confidence level to weigh the predicted price.Subsequently, we summarize the transactions into 8 trading operations, input the modified price and automatically compare the hypothetical return of these eight operations.Then, output the operation with largest return as today's decision.We compare the returns and transaction costs of comparative systems, and demonstrate our strategy with effectiveness.
Xinjia Xie, Shun Gai, Yunxiao Guo, Han Long
SEKE2