EDBT 2026 Demo / reviewers in the wild / expert
Kuankuan Guo
dblp:303/8030
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2024
0009-0007-4202-8537ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SEDIT: Space-Efficient Discriminative Bit Tree for Hybrid Memory Indexing
Yuanjin Lin, Kaixin Huang, Kuankuan Guo, Linpeng Huang |
DASFAA (1) | 4 |
| 2024 | High-Performance Remote Data Persisting for Key-Value Stores via Persistent Memory RegionabstractKey-value stores (KVStores), such as LevelDB and Redis, have been widely used in real-world production environments. To guarantee data durability and availability, traditional KVStores suffer from high write latency, mainly caused by the long network and data-persisting time. To solve this problem, this article presents a novel data-persisting path for KVStores, allowing remote clients to persist data to the KVStore server with$\mu s$-level latency. The novelty of this study is threefold. First, we propose PMRDirect, which utilizes a persistent memory region (PMR) in the NVM express standard to construct a direct data-persisting path from the RDMA networking card (NIC) to the PMR region inside an SSD. Second, to showcase PMRDirect in KVStores, we developed a new accessing stack called PMRAccess, enabling remote clients to access existing KVStores and providing durability for each write request. Specifically, we present a low-latency RDMA-based messaging mode and a chunk-based PMR management in PMRAccess to reduce write latency and improve system throughput. Finally, we conducted extensive experiments to evaluate the performance of our proposals. We first compared PMRDirect with a few remote data-persisting paths to show its effectiveness. Then, we evaluated PMRAccess upon two KVStores, including LibCuckoo (an in-memory KVStore) and LevelDB (an in-storage KVStore). The results showed that PMRAccess outperformed the SSD-based accessing stack by up to$6.1\times $in write throughput and$36\times $in write tail latency, and it achieved$1.7\times $higher write throughput and$0.59\times $lower write tail latency over the PMEM-based accessing stack. Further, we conducted a system-to-system comparison between the PMRAccess-integrated LibCuckoo and Redis, and the results showed our proposal achieved up to$13\times $higher throughputs and$40\times $lower write latency than Redis. Yongping Luo, Peiquan Jin, Zhaole Chu, Kuankuan Guo, Jinhui Guo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | ZoneKV: A Space-Efficient Key-Value Store for ZNS SSDsabstractIn this paper, we propose a new space-efficient key-value store called ZoneKV for ZNS (Zoned Namespace) SSDs. We observe that existing work on adapting RocksDB to ZNS SSDs will cause fragmentation of zones and severe space amplification. Thus, we propose a lifetime-based zone storage model and a level-specific zone allocation algorithm to store SSTables with a similar lifetime in the same zone. We evaluate ZoneKV on a real ZNS SSD. The results show that ZoneKV can reduce up to 60% space amplification and maintain higher throughputs than two competitors, RocksDB and ZenFS. Mingchen Lu, Peiquan Jin, Yongping Luo, Kuankuan Guo |
DAC | 5 |
| 2023 | HM2: Efficient Host Memory Management for RDMA-Enabled Distributed SystemsabstractRemote direct memory access (RDMA) supports zero-copy networking by transferring data from clients directly to host memory, eliminating the need to copy data between clients' memory and the data buffers in the hosting server. However, the hosting server must design efficient memory management schemes to handle incoming clients' data. In this paper, we propose a high-performance host memory management scheme called HM2 for RDMA-enabled distributed systems. We present a new buffer structure for incoming data from clients. In addition, we propose efficient data processing methods to reduce network transfers between clients and servers. We conducted a preliminary experiment to evaluate HM2, and the results showed HM2 achieved higher throughput than existing schemes, including L5 and FaRM. Zhaole Chu, Peiquan Jin, Yongping Luo, Kuankuan Guo |
HPDC | 5 |
| 2023 | Krypton: Real-time Serving and Analytical SQL Engine at ByteDanceabstractIn recent years, at ByteDance, we have started seeing more and more business scenarios that require performing real-time data serving besides complex Ad Hoc analysis over large amounts of freshly imported data. The serving workload requires performing complex queries over massive newly added data items with minimal delay. These systems are often used in mission-critical scenarios, whereas traditional OLAP systems cannot handle such use cases. To work around the problem, ByteDance products often have to use multiple systems together in production, forcing the same data to be ETLed into multiple systems, causing data consistency problems, wasting resources, and increasing learning and maintenance costs. To solve the above problem, we built a single Hybrid Serving and Analytical Processing (HSAP) system to handle both workload types. HSAP is still in its early stage, and very few systems are yet on the market. This paper demonstrates how to build Krypton, a competitive cloud-native HSAP system that provides both excellent elasticity and query performance by utilizing many previously known query processing techniques, a hierarchical cache with persistent memory, and a native columnar storage format. Krypton can support high data freshness, high data ingestion rates, and strong data consistency. We also discuss lessons and best practices we learned in developing and operating Krypton in production. Jianjun Chen 0001, Li Zhang 0132, Liya Fan, Mu Xiong, Benchao Dong, Kuankuan Guo, Yuanjin Lin, Zikang Wang, Yemeng Yang, Junda Zhao, Dongyan Zhou, Zhikai Zuo, Yuming Liang |
Proc. VLDB Endow. | 12 |
| 2023 | Cooperative Buffer Management With Fine-Grained Data Migrations for Hybrid Memory SystemsabstractHybrid memory composed of DRAM and persistent memory (PM) offers a promising way to realize large-capacity main memory supporting in-memory data storage and computing. However, traditional buffer management schemes focus on improving the hit ratio but lack awareness of the limitations of PM, e.g., slower write time and lower write endurance than DRAM. Therefore, developing new buffer management policies that can reduce costly write-backs of PM blocks while maintaining high performance for the hybrid buffer, is of paramount importance. Existing approaches mainly use a page-grained buffering policy, which will cause unnecessary data migrations between DRAM and PM, leading to a high number of disk I/Os and PM writes. Aiming to reduce I/O costs and PM writes, we propose a new buffer manager named HiBuffer for DRAM/PM-based hybrid memory systems. HiBuffer presents several novel ideas. First, it adopts multigrained data layouts to manage the hybrid buffer cooperatively. In addition to the page granularity, we introduce Lines for the DRAM buffer and Sectors to the PM buffer, forming a buffer with three granularities, including Line, Sector, and Page. We prove that the multigrained cooperative buffer management can deliver higher performance than existing page-grained schemes. Second, we propose a sector-grained method to migrate data from DRAM to PM, which can avoid unnecessary data movements and reduce PM writes. Third, we use an out-of-place updating mechanism to absorb updates in DRAM, which can further reduce the writes to PM. We compare HiBuffer with three existing schemes, including LRU, CLOCK-DWF, and MiniPage, on five synthetic workloads and the YCSB benchmark using real Intel Optane DC PM. The results in terms of various metrics, including running time, PM writes, hit ratio, and disk I/Os, suggest the efficiency of HiBuffer. In particular, HiBuffer reduces the running time by up to 37.8% and the writes to PM by up to 83% compared to the competitors when evaluated on the YCSB benchmark. Peiquan Jin, Yongping Luo, Zhaole Chu, Yigui Yuan, Xujian Zhao, Yuanjing Lin, Kuankuan Guo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2022 | ZonedStore: A Concurrent ZNS-Aware Cache System for Cloud Data StorageabstractCloud data storage relies on efficient cache systems to offer high performance for intensive reads/writes on big data. Due to the large data volume of cloud data storage and the limited capacity of DRAM, current cloud vendors prefer to use SSDs (Solid State Drives) but not DRAM to build the cache system. However, traditional SSDs have a serious over-provisioning problem and a high cost in garbage collection. Thus, the performance of SSD-based cache systems will drop quickly when the usage of SSDs increases. Recently, Zoned Namespaces (ZNS) SSDs have emerged as a hot topic in both academics and industries. Compared to conventional SSDs, ZNS SSDs have the advantages of less overhead of garbage collection and lower over-provisioning costs. Therefore, ZNS SSDs have been a better candidate for the cache system for cloud storage. However, ZNS SSDs only accept sequential writes, and the zones inside ZNS SSDs need to be carefully managed to maximize the advantages of ZNS SSDs. Therefore, making the cache system adapt to ZNS SSDs is becoming a challenging issue. In this paper, we demonstrate ZonedStore, a novel ZNS-aware cache system for cloud data storage. After a brief introduction to the architecture of ZonedStore, we present the key designs of ZonedStore, including a Zone Manager to control the space allocation and operations on ZNS SSDs, a Multi-Layer Buffer Manager, and an In-Memory Concurrent Index to accelerate accesses. Finally, we present a case study to demonstrate the working process and performance of ZonedStore. Yanqi Lv 0001, Peiquan Jin, Ruicheng Liu, Yuanjin Lin, Kuankuan Guo |
ICDCS | 7 |
| 2022 | Efficient Data Placement for Zoned Namespaces (ZNS) SSDs
Peiquan Jin, Mingchen Lu, Xiangyu Zhuang, Yuanjing Lin, Kuankuan Guo |
NPC | 7 |
| 2021 | Elastic and Stable Compaction for LSM-tree: A FaaS-Based Approach on TerarkDBabstractLSM-tree is widely used as a write-optimized storage engine in many NoSQL systems. However, the periodical compaction operations in LSM-tree cost many I/O bandwidths and CPU resources of the local server, resulting in throughput drops of the system. To address this issue, this paper proposes a new compaction scheme based on the FaaS (Functions as a Service) architecture, which is called FaaS Compaction. It utilizes the elastic computing capability of FaaS and always pushes compactions to a FaaS cluster. The FaaS cluster will perform actual compaction operations, which will not affect the processing of the local server. Therefore, we can maintain stable performance even when periodical compactions are triggered. We also present a Parallel Slight Compaction method to solve the timeout problem caused by heavy compactions. We implement the FaaS Compaction based on TerarkDB and a real FaaS cluster and experimentally compare the FaaS Compaction with the RocksDB's local compaction scheme and the state-of-the-art offloading compaction policy. The results suggest the efficiency, stability, and elasticity of our proposal. Jianchuan Li, Peiquan Jin, Yuanjin Lin, Kuankuan Guo |
CIKM | 6 |
| 2021 | Supporting Elastic Compaction of LSM-tree with a FaaS ClusterabstractLSM-tree is widely used as a write-optimized storage engine in many key-value stores. However, the periodical compaction operations in LSM-tree cost many I/O bandwidths and CPU resources of the local server, resulting in throughput drops of the system. To address this issue, this paper proposes a new compaction scheme based on a FaaS (Functions as a Service) cluster, which is called FaaS Compaction. It utilizes the elastic computing capability of FaaS clusters and pushes compactions to a FaaS cluster. The FaaS cluster will perform actual compaction operations, which will not affect the processing of the local server. Therefore, we can maintain stable performance even when periodical compactions are triggered. We implement the FaaS Compaction and compare the FaaS Compaction with RocksDB and the state-of-the-art offloading compaction policy. The results suggest the efficiency and elasticity of our proposal. Jianchuan Li, Peiquan Jin, Kuankuan Guo, Yuanjin Lin |
CLUSTER | 4 |