VLDB 2026 Research / reviewers in the wild / expert
Huatao Wu
dblp:362/4234
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0000-7971-5014ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast Cloud Storage for AI Jobs via Grouped I/O API with Transparent Read/Write Optimizations
Yingyi Hao, Ting Yao 0001, Xingda Wei, Dingyan Zhang, Tianle Sun, Zhiyong Fu, Huatao Wu, Rong Chen 0001 |
FAST | 8 |
| 2026 | Cost-efficient Archive Cloud Storage with Tape: Design and Deployment
Qing Wang 0031, Fan Yang 0134, Qiang Liu 0011, Geng Xiao, Yongpeng Chen, Leiming Chen, Bangzhu Chen, Chenrui Liu, Pingchang Bai, Zigan Luo, Mingyu Xie, Yu Wang 0002, Youyou Lu, Huatao Wu, Jiwu Shu |
FAST | 16 |
| 2025 | DShuffle: DPU-Optimized Shuffle Framework for Large-scale Data Processing
Chen Ding 0012, Sicen Li, Kai Lu 0002, Ting Yao 0001, Daohui Wang, Huatao Wu, Jiguang Wan 0001, Zhihu Tan, Changsheng Xie 0001 |
USENIX ATC | 6 |
| 2025 | DFlush: DPU-Offloaded Flush for Disaggregated LSM-based Key-Value StoresabstractRapid increase of storage and network bandwidth incurs higher CPU consumption in modern data systems. This phenomenon is particularly evident for log-structured merged key-value stores (LSM-KVS), which rely on resource-intensive background operations to flush and compact disk data. While extensive research has been conducted to reduce the CPU overhead of background compaction, less attention has been paid to background flushing, which can also consume a significant amount of valuable CPU cycles and disrupt CPU caches, ultimately impacting overall performance. In this paper, we propose DFlush, a novel solution that uses DPUs to offload background flush operations to reduce its CPU cost. DPUs are an appealing choice for this goal due to their cost-effectiveness, ease of programming, and widespread deployment. However, their complex hardware architecture requires careful design of both the data and control planes. To fully harness the DPU's capabilities, DFlush decomposes a flush job into fine-grained steps, mapped them to DPU hardware units, and accelerates them through pipeline, data, and channel parallelism, ensuring data-plane efficiency. It also introduces an adaptive control plane that dynamically schedules flush jobs from different LSM-KVS instances based on their priority, reducing write stall and tail latency. Our experiments on a real DPU platform with an industrial-grade LSM-KVS show that DFlush delivers higher throughput, significantly lower tail latency, and saves up to dozens of CPU cores per LSM-KVS server while reducing energy consumption. Chen Ding 0012, Kai Lu 0002, Quanyi Zhang, Zekun Ye, Ting Yao 0001, Daohui Wang, Huatao Wu, Jiguang Wan 0001 |
Proc. ACM Manag. Data | 7 |
| 2024 | SepHash: A Write-Optimized Hash Index On Disaggregated Memory via Separate Segment StructureabstractDisaggregated memory separates compute and memory resources into independent pools connected by fast RDMA (Remote Direct Memory Access) networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. Hash indexes provide high-performance single-point operations and are widely used in distributed systems and databases. However, under disaggregated memory, existing hash indexes suffer from write performance degradation due to high resize overhead and concurrency control overhead. Traditional write-optimized hash indexes are not efficient for disaggregated memory and sacrifice read performance. In this paper, we propose SepHash, a write-optimized hash index for disaggregated memory. First, SepHash proposes a two-level separate segment structure that significantly reduces the bandwidth consumption of resize operations. Second, SepHash employs a low-latency concurrency control strategy to eliminate unnecessary mutual exclusion and check overhead during insert operations. Finally, SepHash designs an efficient cache and filter to accelerate read operations. The evaluation results show that, compared to state-of-the-art distributed hash indexes, SepHash achieves a 3.3X higher write performance while maintaining comparable read performance. Xinhao Min, Kai Lu 0002, Jiguang Wan 0001, Changsheng Xie 0001, Daohui Wang, Ting Yao 0001, Huatao Wu |
Proc. VLDB Endow. | 8 |
| 2024 | Scythe: A Low-latency RDMA-enabled Distributed Transaction System for Disaggregated MemoryabstractDisaggregated memory separates compute and memory resources into independent pools connected by RDMA (Remote Direct Memory Access) networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. However, existing RDMA-based distributed transactions on disaggregated memory suffer from severe long-tail latency under high-contention workloads. In this article, we propose Scythe, a novel low-latency RDMA-enabled distributed transaction system for disaggregated memory. Scythe optimizes the latency of high-contention transactions in three approaches: (1) Scythe proposes a hot-aware concurrency control policy that uses optimistic concurrency control (OCC) to improve transaction processing efficiency in low-conflict scenarios. Under high conflicts, Scythe designs a timestamp-ordered OCC (TOCC) strategy based on fair locking to reduce the number of retries and cross-node communication overhead. (2) Scythe presents an RDMA-friendly timestamp service for improved timestamp management. And, (3) Scythe designs an RDMA-optimized RPC framework to improve RDMA bandwidth utilization. The evaluation results show that, compared with state-of-the-art distributed transaction systems, Scythe achieves more than 2.5× lower latency with 1.8× higher throughput under high-contention workloads. Kai Lu 0002, Siqi Zhao, Haikang Shan, Guokuan Li, Jiguang Wan 0001, Ting Yao 0001, Huatao Wu, Daohui Wang |
ACM Trans. Archit. Code Optim. | 8 |
| 2024 | Rcmp: Reconstructing RDMA-Based Memory Disaggregation via CXLabstractMemory disaggregation is a promising architecture for modern datacenters that separates compute and memory resources into independent pools connected by ultra-fast networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. However, existing memory disaggregation solutions based on remote direct memory access (RDMA) suffer from high latency and additional overheads including page faults and code refactoring. Emerging cache-coherent interconnects such as CXL offer opportunities to reconstruct high-performance memory disaggregation. However, existing CXL-based approaches have physical distance limitation and cannot be deployed across racks. In this article, we propose Rcmp, a novel low-latency and highly scalable memory disaggregation system based on RDMA and CXL. The significant feature is that Rcmp improves the performance of RDMA-based systems via CXL, and leverages RDMA to overcome CXL’s distance limitation. To address the challenges of the mismatch between RDMA and CXL in terms of granularity, communication, and performance, Rcmp (1) provides a global page-based memory space management and enables fine-grained data access, (2) designs an efficient communication mechanism to avoid communication blocking issues, (3) proposes a hot-page identification and swapping strategy to reduce RDMA communications, and (4) designs an RDMA-optimized RPC framework to accelerate RDMA transfers. We implement a prototype of Rcmp and evaluate its performance by using micro-benchmarks and running a key-value store with YCSB benchmarks. The results show that Rcmp can achieve 5.2× lower latency and 3.8× higher throughput than RDMA-based systems. We also demonstrate that Rcmp can scale well with the increasing number of nodes without compromising performance. Zhonghua Wang 0001, Yixing Guo, Kai Lu 0002, Jiguang Wan 0001, Daohui Wang, Ting Yao 0001, Huatao Wu |
ACM Trans. Archit. Code Optim. | 7 |
| 2023 | DoW-KV: A DPU-offloaded and Write-optimized Key-Value Store on Disaggregated Persistent MemoryabstractDisaggregated Persistent Memory (DPM) is a promising technology offering elasticity, high resource utilization, persistent data storage, and lower power consumption. While building KV stores on the DPM benefits from these merits, achieving efficient writes also faces two primary challenges: 1) limited scalability caused by the underused PM bandwidth, and 2) limited CPU resources the persistent memory server (PMS) can provide. Integrating the SmartNIC such as the Data Processing Unit (DPU) into the DPM gives developers the chance to optimize writing to KV stores by utilizing both the memory and processor of DPU. However, simple offloading cannot make full use of the DPU’s potential capacity. To address these challenges, we propose DoW-KV, a persistent hash KV store on DPM. DoW-KV employs a two-tier hash index consisting of a DPU cache table in DPU memory and multiple PM persistent tables on the PM. It relocates small random writes to the DPU memory and consolidates them to the PM at a coarse granularity. Furthermore, DoW-KV uses DPU-offloaded step merge and a coroutine-based asynchronous processing framework to efficiently manage the PM persistent tables. DoW-KV also introduces a client-mixed read strategy to boost key searching on the two-tier hash index. Experimental results show that DoW-KV outperforms the state-of-the-art DINOMO by 2.1× and 1.3× in the Put and Get operations, respectively. Guokuan Li, Jiguang Wan 0001, Junyue Wang, Ting Yao 0001, Huatao Wu, Daohui Wang |
CLUSTER | 7 |