VLDB 2026 Research / reviewers in the wild / expert
Keji Huang
dblp:298/5291
· DBLP profile ↗
13ranked-venue papers
0as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Achieving Availability, Efficiency, and Elasticity in Stripeless Erasure-Coded StorageabstractErasure coding plays a crucial role in distributed storage systems to provide fault tolerance at a low storage cost. Conventional erasure coding schemes determine data placement based on stripes. However, placing data into stripes can incur non-negligible performance overheads that will manifest in emerging fast in-memory storage systems, making conventional erasure coding schemes suboptimal in such scenarios. Aiming to eliminate such overheads, we present Nos , a stripeless placement scheme for erasure-coded distributed in-memory storage. It lets each node independently replicate data to other nodes and encode received data replicas into parities with XOR. Thus, it avoids the overheads caused by stripes. To enable failure recovery, Nos uses a combinatoric structure called symmetric balanced incomplete block design (SBIBD) to decide primary-to-backup node affinities during replication. Atop Nos , we further build Nostor , a distributed in-memory key-value store. We also achieve high availability and efficient foreground serving simultaneously with Hotness-aware Two-phase Reconstruction ( HR ). Evaluations demonstrate that Nostor achieves 1.61× to 2.60× throughputs compared to Cocytus , PQ , and Split with similar or lower latencies than these stripe-based erasure coding baselines. Equipped with HR , Nostor + HR also achieves 46.7% P999 foreground latency reduction during node repair. Zeqi Li, Zhirong Shen, Yuhao Zhang 0006, Keji Huang, Jiwu Shu |
ACM Trans. Storage | 6 |
| 2026 | HyTorC: Hybrid Address Translation for SSDs Supporting CompressionabstractHigh-capacity solid-state drives (SSDs) with expected capacities of one PByte and more will address cloud storage and archiving environments previously dominated by magnetic disks. These new applications are very cost-sensitive, so unnecessary overhead must be reduced as much as possible without sacrificing the performance advantages of SSDs. An expensive component within a scale-up SSD is the on-device memory to map host logical page numbers to flash pages. This article, therefore, proposes HyTorC, which builds on the idea of S-FTL to represent sequentially stored logical pages using bitmaps instead of providing one entry per logical page. HyTorC extends this idea by introducing, for the first time in an FTL, compression of these bitmaps using run-length encoding, and by investigating the effects of background scrubbing to realign randomly written pages into contiguous runs. This background scrubbing allows HyTorC to keep large portions of the mapping table in block mapping mode, further reducing the memory footprint. HyTorC retains the flexibility of the page mapping scheme and supports compression of logical blocks within the SSD, allowing multiple compressed logical pages to be stored within a single physical page. HyTorC is fully implemented in an open-channel SSD. Our tests show that HyTorC can reduce memory consumption by an average of 98.9% over standard page mapping, 95.6% over the DFTL scheme, 87.3% over S-FTL, and 56.5% over the learned index-based approach LeaFTL for the Alibaba Cloud block traces, and by 98.6% over standard page mapping, 95.5% over the DFTL scheme, 84.1% over S-FTL, and 61.5% over LeaFTL for the Microsoft Research Cambridge traces. HyTorC focuses on the memory footprint of the FTL and not on performance. However, the performance evaluation shows that HyTorC achieves similar performance compared to page mapping and FTLs based on learned indexes. Yu Zhang 0294, Renhai Chen, Gong Zhang 0001, Peng Wang 0037, Xin Yao 0008, Keji Huang, André Brinkmann |
ACM Trans. Storage | 6 |
| 2025 | Stripeless Data Placement for Erasure-Coded In-Memory Storage
Jiwu Shu, Yuhao Zhang 0006, Keji Huang |
OSDI | 5 |
| 2025 | HLN-Tree: A memory-efficient B+-Tree with huge leaf nodes and locality predictorsabstractKey-value stores in Cloud environments can contain more than 2 45 unique elements and be larger than 100 PByte. B + -Trees are well suited for these larger-than-memory datasets and seamlessly index data stored on thousands of secondary storage devices. Unfortunately, it is often uneconomical to even store all inner tree nodes in memory for these dataset sizes. Therefore, lookup performance is affected by the additional IOs for reading inner nodes. This number of inner nodes can be reduced by increasing the size of leaf nodes. We propose HLN-Trees, which support huge leaf nodes without increasing the IO sizes for individual index operations. They partition leaf nodes in arrays of independent subnodes and combine ideas from BD-trees with rebalancing, learning key deviations, and storing locality predictors. HLN-Trees have been initially designed for uniform random key distributions and support arbitrary key distributions through an additional layer of hashing in leaf nodes. HLN-Trees decrease the number of inner nodes by up to 256× for uniform random key distributions and by 16× to 64× for arbitrary ones compared to B + -Trees, while keeping their performance at the same level even at high concurrency levels. We show analytically and through real-world and synthetic benchmarks that HLN-Trees also outperform state-of-the-art learned indexes for secondary storage. André Brinkmann, Reza Salkhordeh, Florian Wiegert, Peng Wang 0037, Xin Yao 0008, Renhai Chen, Keji Huang, Gong Zhang 0001 |
ACM Trans. Storage | 7 |
| 2025 | Efficiently Enlarging RDMA-Attached Memory with SSDabstractRDMA-based in-memory storage systems offer high performance but are restricted by the capacity of physical memory. In this article, we propose TeRM to extend RDMA-attached memory with SSD. TeRM achieves fast remote access on the SSD-extended memory by eliminating page faults of RDMA NIC and CPU from the critical path. We also introduce a set of techniques to reduce the consumption of CPU and network resources. Evaluation shows that TeRM performs close to the performance of the ideal upper bound where all pages are pinned in the physical memory. Compared with existing approaches, TeRM significantly improves the performance of unmodified RDMA-based storage systems, including a file system and a key-value system. Zhe Yang 0012, Qing Wang 0031, Xiaojian Liao, Youyou Lu, Keji Huang, Jiwu Shu |
ACM Trans. Storage | 5 |
| 2024 | TeRM: Extending RDMA-Attached Memory with SSD
Zhe Yang 0012, Qing Wang 0031, Xiaojian Liao, Youyou Lu, Keji Huang, Jiwu Shu |
FAST | 5 |
| 2024 | Separation Is for Better Reunion: Data Lake Storage at HuaweiabstractHuawei collaborates with some Chinese large busi-ness companies to store and process exabytes of nationwide operational data in data lake storage to provide business insights. Specifically, our customers will ask to store and process massive log message data to support their real-time and decision-making applications. Thus, we need computation and storage components in the analytic platform to process and store these data cost-efficiently. To meet these user requirements, we have designed a storage system in data lake, StreamLake, which introduces a novel design to serve log message streaming and batch data processing in distributed storage, with high scalability, efficiency, reliability and low cost. Specifically, we introduce a stream (storage) object as a storage abstraction for message streaming data to achieve the storage-disaggregated architecture with high scalability and reliability. Moreover, we utilize the erasure coding and tiered storage to save the storage cost, and furthermore, the stream object can be automatically converted to a table object such that cost-effective stream and batch data processing can be achieved. For tabular data, we implement the lakehouse functionality to support ACID via the table object, with a metadata acceleration to improve the efficiency of data access between the compute and storage engines. Also, we design a LakeBrain optimizer at the storage side to optimize the query performance and resource utilization under the storage-disaggregated architecture. Finally, we have also deployed StreamLake in China Mobile, the world's largest mobile network operator to serve over 20PB production data, and the results demonstrate improvements of 30% to 4x in terms of query performance and over 37% in terms of cost saving. Chengliang Chai, Haohai Ma, Zhenyong Fan, Jiaquan Zhang, Rui Zhang 0003, Duanshun Li, Keji Huang, Guangbin Meng, Yuefeng Zhou, Lirong Jian, Jiwu Shu, Ye Yuan 0001, Guoren Wang, Guoliang Li 0001 |
ICDE | 12 |
| 2024 | Composable Storage Servers: A Storage Paradigm for Disaggregated SystemsabstractDisaggregated and composable data centers optimize resource allocation and mitigate overprovisioning through dynamic resource management. While significant research has concentrated on disaggregating components and composing compute servers, the application of these principles to storage servers within disaggregated data centers remains underexplored. Traditional storage servers are often overprovisioned to accommodate a wide range of scenarios and workloads, resulting in designs that conflict with the principles of composable data centers, which prioritize efficiency by minimizing resource over-provisioning across components. This paper presents the concept of Composable Storage Servers, a design approach for storage solutions within disaggregated data centers. By applying the principles of resource disaggregation, this approach mitigates the overprovisioning of resources in storage servers. Central to this vision is the Core Storage Node, which integrates essential storage functionalities into a unified component, consistent with composable infrastructure principles. Our prototype and real-world deployment of the core storage node show its effectiveness at substituting local SSDs, achieving a 30% reduction in storage costs. Shai Bergman, Onur Mutlu, Wu Yong, Keji Huang, Ji Zhang 0035 |
NAS | 4 |
| 2024 | Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory AccessabstractThe growing memory demands of modern applications have driven the adoption of far memory technologies in data centers to provide cost-effective, high-capacity memory solutions. However, far memory presents new performance challenges because its access latencies are significantly longer and more variable than local DRAM. For applications to achieve acceptable performance on far memory, a high degree of memory-level parallelism (MLP) is needed to tolerate the long access latency. While modern out-of-order processors are capable of exploiting a certain degree of MLP, they are constrained by resource limitations and hardware complexity. The key obstacle is the synchronous memory access semantics of traditional load/store instructions, which occupy critical hardware resources for a long time. The longer far memory latencies exacerbate this limitation. This article proposes a set of Asynchronous Memory Access Instructions (AMI) and its supporting function unit, Asynchronous Memory Access Unit (AMU), inside contemporary Out-of-Order Core. AMI separates memory request issuing from response handling to reduce resource occupation. Additionally, AMU architecture supports up to several hundreds of asynchronous memory requests through re-purposing a portion of L2 Cache as scratchpad memory (SPM) to provide sufficient temporal storage. Together with a coroutine-based programming framework, this scheme can achieve significantly higher MLP for hiding far memory latencies. Evaluation with a cycle-accurate simulation shows AMI achieves 2.42× speedup on average for memory-bound benchmarks with 1μs additional far memory latency. Over 130 outstanding requests are supported with 26.86× speedup for GUPS (random access) with 5 μs latency. These demonstrate how the techniques tackle far memory performance impacts through explicit MLP expression and latency adaptation. Luming Wang, Xu Zhang 0033, Songyue Wang, Zhuolun Jiang, Tianyue Lu, Mingyu Chen 0001, Siwei Luo, Keji Huang |
ACM Trans. Archit. Code Optim. | 8 |
| 2023 | Revisiting Swapping in User-Space With Lightweight ThreadingabstractMemory-intensive applications, such as in-memory databases, caching systems, and key-value stores, are increasingly demanding larger main memory to fit their working sets. Conventional swapping can enlarge the memory capacity by paging out inactive pages to backend stores. However, existing swapping solutions suffer several performance and compatibility issues, making them unsuitable for high-concurrency and memory-intensive applications. In this article, we redesign the swapping system and propose Lightswap, a high-performance user-space swapping solution that supports paging with both local SSDs and remote memories. First, to avoid kernel involvement, we propose to leverage the extended Berkeley packet filter (eBPF) for handling page faults (PFs) in user space and further eliminate the heavy I/O stack with the help of user-space I/O drivers. Then, we co-design the PF handling with lightweight thread (LWT) scheduling to improve system throughput and reduce the end-to-end PF latency. Finally, we propose a try-catch framework in Lightswap to deal with swap-in errors which have been exacerbated by the scaling in process technology. We implement Lightswap in our production-level system and evaluate it with various benchmarks. Results show that Lightswap achieves scalable PF notification latency ($4 \mu \text{s}$under 128 LWTs), reduces the PF handling latency by 3–5 times, and improves the throughput of memcached by more than 40% compared with the state-of-art swapping systems. Kan Zhong, Wenlin Cui, Qiao Li 0001, Zhe Yang 0012, Youyou Lu, Xiaodan Yan, Siwei Luo, Qizhao Yuan, Keji Huang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 10 |
| 2022 | L-QoCo: learning to optimize cache capacity overloading in storage systemsabstractCache plays an important role to maintain high and stable performance (i.e. high throughput, low tail latency and throughput jitter) in storage systems. Existing rule-based cache management methods, coupled with engineers' manual configurations, cannot meet ever-growing requirements of both time-varying workloads and complex storage systems, leading to frequent cache overloading. Xijun Li, Xiyao Zhou, Mingxuan Yuan, Keji Huang |
DAC | 6 |
| 2022 | Pacman: An Efficient Compaction Approach for Log-Structured Key-Value Store on Persistent Memory
Jing Wang 0158, Youyou Lu, Qing Wang 0031, Minhui Xie, Keji Huang, Jiwu Shu |
USENIX ATC | 5 |
| 2022 | SwitchTx: Scalable In-Network Coordination for Distributed Transaction ProcessingabstractOnline-transaction-processing (OLTP) applications require the underlying storage system to guarantee consistency and serializability for distributed transactions involving large numbers of servers, which tends to introduce high coordination cost and cause low system performance. In-network coordination is a promising approach to alleviate this problem, which leverages programmable switches to move a piece of coordination functionality into the network. This paper presents a fast and scalable transaction processing system called SwitchTx. At the core of SwitchTx is a decentralized multi-switch in-network coordination mechanism, which leverages modern switches' programmability to reduce coordination cost while avoiding the central-switch-caused problems in the state-of-the-art Eris transaction processing system. SwitchTx abstracts various coordination tasks (e.g., locking, validating, and replicating) as in-switch gather-and-scatter (GaS) operations, and offloads coordination to a tree of switches for each transaction (instead of to a central switch for all transactions) where the client and the participants connect to the leaves. Moreover, to control the transaction traffic intelligently, SwitchTx reorders the coordination messages according to their semantics and redesigns the congestion control combined with admission control. Evaluation shows that SwitchTx outperforms current transaction processing systems in various workloads by up to 2.16X in throughput, 40.4% in latency, and 41.5% in lock time. Youyou Lu, Yiming Zhang 0003, Qing Wang 0031, Keji Huang, Jiwu Shu |
Proc. VLDB Endow. | 6 |