Fan Guo 0003

dblp:55/6662-3 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
5since 2021 · last 2025
0009-0007-6842-5899ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Split-LSM-Tree: High-Performance KV Storage via Dynamic Tree Splitting
abstract
Many key-value storage systems employ LSM-tree-based architectures, which often suffer from significant read and write amplification issues that can substantially impact overall performance. While existing solutions mitigate amplification through key-value separation or relaxed ordering, they retain deep hierarchical structures that perpetuate scalability bottlenecks. We present Split-LSM-Tree, a novel storage engine that dynamically partitions monolithic LSM-trees into a forest of capacity-bounded subtrees when thresholds are breached. Our design introduces three key innovations: (1) An adaptive splitting trigger mechanism that initiates partitioning based on storage pressure and performance degradation; (2) A non-blocking splitting algorithm enabling seamless subtree migration; and (3) I/O optimizations preserving foreground operation continuity. Our experiments show that compared to the widely deployed LevelDB and RocksDB, Split-LSM-Tree reduces write amplification by$\mathbf{2 5. 7} \boldsymbol{\%}$and$\mathbf{1 8. 1 \%}$respectively, while achieving$\mathbf{2. 9} \times$and$\mathbf{2. 4} \times$higher write throughput.
Yu-Ang Cao, Fan Guo 0003, Deming Ren, Ruida Xu, Wenzhe Zhu, Yongkun Li 0001, Yinlong Xu 0001
ICPADS2
2025 Double-Metadata-Cache: An Efficient Metadata Caching Architecture for Distributed File Systems
abstract
Current distributed file systems manage metadata through a flattened metadata service, which requires accessing metadata across multiple metadata servers. To improve performance, file systems use client-side cache to reduce frequent server access and employ a lease mechanism to maintain client metadata cache consistency. However, the lease mechanism blocks metadata updates and requires periodic renewal of cache entries. To address these challenges, we propose Double-Metadata-Cache, an efficient two-layer metadata caching architecture. This solution incorporates two key designs: a validity period-based lazy update mechanism to maintain client metadata cache consistency, a segmented caching strategy uses hot prefixes to reduce accesses to metadata servers. Experimental results show that Double-Metadata-Cache achieves higher performance in high-proportion metadata operations such as creation and stat. Moreover, as the file depth increases, it reduces path resolution latency by 50 % compared to the client-side caching strategy.
Fan Guo 0003, Yu-Ang Cao, Deming Ren, Wenzhe Zhu, Yongkun Li 0001, Yinlong Xu 0001
ICPADS2
2025 Amber: Towards Fast and Space-Efficient Incremental Checkpointing in Large Language Model Training
abstract
Large Language Models have transformed numerous domains with their exceptional capabilities. However, their training processes are inherently prone to failures due to the massive scale of training clusters, involving thousands to tens of thousands of GPUs and requiring several months of uninterrupted computation. Checkpointing serves as a critical mechanism for ensuring fault tolerance. However, the substantial size of LLM checkpoints, attributed to their vast number of parameters, poses significant challenges in terms of time and storage overheads. While existing approaches offer partial solutions, they remain inadequate in addressing the unique demands and complexities of LLM training at scale. To this end, we propose Amber, a novel LLM training framework that significantly improves checkpointing speed and storage efficiency by leveraging selective incremental checkpointing, a technique that selectively checkpoints significant parameter updates while omitting minor changes. We implement Amber atop PyTorch, a widely adopted deep learning framework, and demonstrate that Amber achieves up to 10–155 × faster checkpointing compared to state-of-the-art full checkpointing methods across models of varying scales on a single GPU, while maintaining storage overhead below 3%.
Wenzhe Zhu, Chaomei Yan, Fan Guo 0003, Yongkun Li 0001, Yinlong Xu 0001
ICPP5
2023 Towards High Performance and Efficient Memory Deduplication via Mixed Pages
abstract
Large pages are widely supported in modern hardware and OSes to reduce the overhead of TLB misses. However, memory deduplication can be inefficient with large pages, leading to low memory utilization. To simultaneously enjoy the benefits of high performance by accessing memory with large pages (e.g., 2 MB pages) and high deduplication rate by managing memory with base pages (e.g., 4 KB pages), we proposeSmartMemoryDeduplciation (SmartMD), which is an adaptive and efficient memory management scheme via mixed pages. Specifically, we propose lightweight schemes to periodically monitor pages’ access frequency and repetition rate, and present an adaptive conversion scheme to selectively split or reconstruct large pages. We further optimize SmartMD by developing SmartMD$^{+}$, which dynamically adjusts the page scanning cycle by monitoring the TLB miss cost, and reconstructs the split large pages in an on-demand way so as to reduce the CPU overhead of SmartMD. We further implement a prototype system and conduct extensive experiments with various workloads under different system settings. Experiment results show that SmartMD and SmartMD$^{+}$can simultaneously achieve high access performance similar to systems using large pages, and achieve a deduplication rate similar to that applying aggressive deduplication scheme (i.e., KSM) on base pages.
Lulu Yao, Yongkun Li 0001, Fan Guo 0003, Si Wu 0003, Yinlong Xu 0001, John C. S. Lui
IEEE Trans. Computers3
2023 Dynamic GPU Scheduling With Multi-Resource Awareness and Live Migration Support
abstract
In clouds and data centers, GPU servers with multiple GPUs are widely deployed. Current state-of-the-art GPU scheduling policies are “static” in assigning applications to different GPUs. These policies usually ignore the dynamics of the GPU utilization and are often inaccurate in estimating resource demand before assigning/running applications, so there is a large opportunity to further balance the loads and improve GPU utilization. Based on CUDA (Compute Unified Device Architecture), we develop a runtime system called DCUDA which supports“dynamic”scheduling of running applications between multiple GPUs. In particular, DCUDA takes into consideration multidimensional resources, including computing cores, memory usage, and energy consumption. It first provides a real-time and lightweight method to accurately monitor the resource demand of applications and GPU utilization. Furthermore, it provides a universal migration facility to migrate“running applications”between GPUs with negligible overhead. More importantly, DCUDA transparently supports all CUDA applications without changing their source code. Experiments with our prototype system show that DCUDA can reduce 78.3% of overloaded time of GPUs on average. As a result, for different workloads consisting of a wide range of applications we studied, DCUDA can reduce the average execution time of general applications by up to 42.1%, and even up to 67% for memory-intensive tasks. Besides, DCUDA also reduces 13.3% of energy in light-load scenarios.
Yongkun Li 0001, Fan Guo 0003, Yinlong Xu 0001, John C. S. Lui
IEEE Trans. Cloud Comput.3
2019 DCUDA: Dynamic GPU Scheduling with Live Migration Support
abstract
In clouds and data centers, GPU servers which consist of multiple GPUs are widely deployed. Current state-of-the-art GPU scheduling algorithm are "static" in assigning applications to different GPUs. These algorithms usually ignore the dynamics of the GPU utilization and are often inaccurate in estimating resource demand before assigning/running applications, so there is a large opportunity to further load balance and to improve GPU utilization. Based on CUDA (Compute Unified Device Architecture), we develop a runtime system called DCUDA which supports "dynamic" scheduling of running applications between multiple GPUs. In particular, DCUDA provides a realtime and lightweight method to accurately monitor the resource demand of applications and GPU utilization. Furthermore, it provides a universal migration facility to migrate "running applications" between GPUs with negligible overhead. More importantly, DCUDA transparently supports all CUDA applications without changing their source codes. Experiments with our prototype system show that DCUDA can reduce 78.3% of overloaded time of GPUs on average. As a result, for different workloads consisting of a wide range applications we studied, DCUDA can reduce the average execution time of applications by up to 42.1%. Furthermore, DCUDA also reduces 13.3% energy in the light load scenario.
Fan Guo 0003, Yongkun Li 0001, John C. S. Lui, Yinlong Xu 0001
SoCC1
2019 HP-Mapper: A High Performance Storage Driver for Docker Containers
abstract
Docker containers are widely deployed to provide lightweight virtualization, and they have many desirable features such as ease of deployment and near bare-metal performance. However, both the performance and cache efficiency of containers are still limited by their storage drivers due to the coarse-grained copy-on-write operations, and the large amount of redundancy in both I/O requests and page cache. To improve I/O performance and cache efficiency of containers, we develop HP-Mapper, a high performance storage driver for Docker containers. HP-Mapper provides a two-level mapping strategy to support fine-grained copy-on-write with low overhead, and an efficient interception method to reduce redundant I/Os. Furthermore, it uses a novel cache management mechanism to reduce duplicate cached data. Experiment results with our prototype system show that HP-Mapper significantly reduces copy-on-write latency due to its finer-grained copy-on-write scheme. Moreover, HP-Mapper can also reduce 65.4% cache usage on average due to elimination of duplicated data. As a result, HP-Mapper improves the throughput of real-world workloads by up to 39.4%, and improves the startup speed of containers by 2.0x.
Fan Guo 0003, Yongkun Li 0001, Min Lv, Yinlong Xu 0001, John C. S. Lui
SoCC1
2019 Leveraging Array Mapped Tries in KSM for Lightweight Memory Deduplication
abstract
In cloud computing, how to use limited hardware resources to meet the increasing demands has become a major issue. KSM (Kernel Same-page Merging) is a content-based page sharing mechanism used in Linux that merges equal memory pages, thereby significantly reducing memory usage and increasing the density of virtual machines or containers. However, KSM introduces a large overhead in CPU and memory bandwidth usage due to the use of red-black trees and content-based page comparison. To reduce the deduplication overhead, in this paper, we propose a new design called AMT-KSM, which leverages array mapped tries to realize lightweight memory deduplication. The basic idea is to divide each memory page into multiple segments and use the concatenated strings of the hash values of segments as indexed keys in the tries. By doing this, we can significantly reduce the time required for searching duplicate pages as well as the number of page comparisons. We conduct experiments to evaluate the performance of our design, and results show that compared with the conventional KSM, AMT-KSM can reduce up to 44.9% CPU usage and 31.6% memory bandwidth usage.
Lingjing You, Yongkun Li 0001, Fan Guo 0003, Yinlong Xu 0001, Jinzhong Chen, Liu Yuan 0002
NAS3
2019 ElasticBF: Elastic Bloom Filter with Hotness Awareness for Boosting Read Performance in Large Key-Value Stores
Yongkun Li 0001, Chengjin Tian, Fan Guo 0003, Cheng Li 0001, Yinlong Xu 0001
USENIX ATC3
2018 ElasticBF: Fine-grained and Elastic Bloom Filter Towards Efficient Read for LSM-tree-based KV Stores
Yueming Zhang, Yongkun Li 0001, Fan Guo 0003, Cheng Li 0001, Yinlong Xu 0001
HotStorage3
2017 SmartMD: A High Performance Deduplication Engine with Mixed Pages
Fan Guo 0003, Yongkun Li 0001, Yinlong Xu 0001, Song Jiang 0001, John C. S. Lui
USENIX ATC1