EDBT 2026 Demo / reviewers in the wild / expert
Daohui Wang
dblp:332/1457
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LithoMamba: High-fidelity lithography simulation with State Space ModelsabstractLithography simulation is a critical technology in modern semiconductor manufacturing, yet existing deep learning models often fail to accurately model the complex, long-range optical physics due to the inherent locality of convolution. This limitation results in insufficient simulation fidelity and poses significant challenges for optimization tasks. To overcome this challenge, we introduce LithoMamba, the first generative framework to leverage Mamba for high-fidelity lithography simulation. Our architecture uses a Mamba Generator to model global and long-range optical interactions, while a local, MLP-free Discriminator provides precise, spatial feedback to ensure fine-grained pattern fidelity. This global-local design enables our model to achieve both physical realism and exceptional detail. Our experiments show that LithoMamba outperforms existing methods, both in quantitative and qualitative results. These findings demonstrate the promise of State Space Models for improving lithography simulation and suggest new possibilities for combining physics with generative AI in chip manufacturing. Daohui Wang, Shujing Lyu, Pourya Shamsolmoali, Jiwei Shen, Yue Lu 0001 |
DATE | 2 |
| 2026 | Feature Enhancement Module Based on Class-Centric Loss for Fine-Grained Visual ClassificationabstractWe propose a novel feature enhancement module designed for fine-grained visual classification tasks, which can be seamlessly integrated into various backbone architectures, including both convolutional neural network (CNN)-based and Transformer-based networks. The plug-and-play module outputs pixel-level feature maps and performs a weighted fusion of filtered features to enhance fine-grained feature representation. We introduce a class-centric loss function that optimizes the alignment of samples with their target class centers by pulling them toward the center of the target class while simultaneously pushing them away from the center of the most visually similar nontarget classes. Soft labels are employed to mitigate overfitting, ensuring the model generalizes well to unseen examples. Our approach consistently delivers significant improvements in accuracy across various mainstream backbone architectures, underscoring its versatility and robustness. Furthermore, we achieved the highest accuracy on the NABirds (NAB) and our proprietary lock cylinder datasets. We have released our source code and pretrained model on GitHub: https://github.com/Richard5413/FEM-CC.git. Daohui Wang, Shujing Lyu, Tian Wei, Yue Lu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | DShuffle: DPU-Optimized Shuffle Framework for Large-scale Data Processing
Chen Ding 0012, Sicen Li, Kai Lu 0002, Ting Yao 0001, Daohui Wang, Huatao Wu, Jiguang Wan 0001, Zhihu Tan, Changsheng Xie 0001 |
USENIX ATC | 5 |
| 2025 | DFlush: DPU-Offloaded Flush for Disaggregated LSM-based Key-Value StoresabstractRapid increase of storage and network bandwidth incurs higher CPU consumption in modern data systems. This phenomenon is particularly evident for log-structured merged key-value stores (LSM-KVS), which rely on resource-intensive background operations to flush and compact disk data. While extensive research has been conducted to reduce the CPU overhead of background compaction, less attention has been paid to background flushing, which can also consume a significant amount of valuable CPU cycles and disrupt CPU caches, ultimately impacting overall performance. In this paper, we propose DFlush, a novel solution that uses DPUs to offload background flush operations to reduce its CPU cost. DPUs are an appealing choice for this goal due to their cost-effectiveness, ease of programming, and widespread deployment. However, their complex hardware architecture requires careful design of both the data and control planes. To fully harness the DPU's capabilities, DFlush decomposes a flush job into fine-grained steps, mapped them to DPU hardware units, and accelerates them through pipeline, data, and channel parallelism, ensuring data-plane efficiency. It also introduces an adaptive control plane that dynamically schedules flush jobs from different LSM-KVS instances based on their priority, reducing write stall and tail latency. Our experiments on a real DPU platform with an industrial-grade LSM-KVS show that DFlush delivers higher throughput, significantly lower tail latency, and saves up to dozens of CPU cores per LSM-KVS server while reducing energy consumption. Chen Ding 0012, Kai Lu 0002, Quanyi Zhang, Zekun Ye, Ting Yao 0001, Daohui Wang, Huatao Wu, Jiguang Wan 0001 |
Proc. ACM Manag. Data | 6 |
| 2024 | SepHash: A Write-Optimized Hash Index On Disaggregated Memory via Separate Segment StructureabstractDisaggregated memory separates compute and memory resources into independent pools connected by fast RDMA (Remote Direct Memory Access) networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. Hash indexes provide high-performance single-point operations and are widely used in distributed systems and databases. However, under disaggregated memory, existing hash indexes suffer from write performance degradation due to high resize overhead and concurrency control overhead. Traditional write-optimized hash indexes are not efficient for disaggregated memory and sacrifice read performance. In this paper, we propose SepHash, a write-optimized hash index for disaggregated memory. First, SepHash proposes a two-level separate segment structure that significantly reduces the bandwidth consumption of resize operations. Second, SepHash employs a low-latency concurrency control strategy to eliminate unnecessary mutual exclusion and check overhead during insert operations. Finally, SepHash designs an efficient cache and filter to accelerate read operations. The evaluation results show that, compared to state-of-the-art distributed hash indexes, SepHash achieves a 3.3X higher write performance while maintaining comparable read performance. Xinhao Min, Kai Lu 0002, Jiguang Wan 0001, Changsheng Xie 0001, Daohui Wang, Ting Yao 0001, Huatao Wu |
Proc. VLDB Endow. | 6 |
| 2024 | Scythe: A Low-latency RDMA-enabled Distributed Transaction System for Disaggregated MemoryabstractDisaggregated memory separates compute and memory resources into independent pools connected by RDMA (Remote Direct Memory Access) networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. However, existing RDMA-based distributed transactions on disaggregated memory suffer from severe long-tail latency under high-contention workloads. In this article, we propose Scythe, a novel low-latency RDMA-enabled distributed transaction system for disaggregated memory. Scythe optimizes the latency of high-contention transactions in three approaches: (1) Scythe proposes a hot-aware concurrency control policy that uses optimistic concurrency control (OCC) to improve transaction processing efficiency in low-conflict scenarios. Under high conflicts, Scythe designs a timestamp-ordered OCC (TOCC) strategy based on fair locking to reduce the number of retries and cross-node communication overhead. (2) Scythe presents an RDMA-friendly timestamp service for improved timestamp management. And, (3) Scythe designs an RDMA-optimized RPC framework to improve RDMA bandwidth utilization. The evaluation results show that, compared with state-of-the-art distributed transaction systems, Scythe achieves more than 2.5× lower latency with 1.8× higher throughput under high-contention workloads. Kai Lu 0002, Siqi Zhao, Haikang Shan, Guokuan Li, Jiguang Wan 0001, Ting Yao 0001, Huatao Wu, Daohui Wang |
ACM Trans. Archit. Code Optim. | 9 |
| 2024 | Rcmp: Reconstructing RDMA-Based Memory Disaggregation via CXLabstractMemory disaggregation is a promising architecture for modern datacenters that separates compute and memory resources into independent pools connected by ultra-fast networks, which can improve memory utilization, reduce cost, and enable elastic scaling of compute and memory resources. However, existing memory disaggregation solutions based on remote direct memory access (RDMA) suffer from high latency and additional overheads including page faults and code refactoring. Emerging cache-coherent interconnects such as CXL offer opportunities to reconstruct high-performance memory disaggregation. However, existing CXL-based approaches have physical distance limitation and cannot be deployed across racks. In this article, we propose Rcmp, a novel low-latency and highly scalable memory disaggregation system based on RDMA and CXL. The significant feature is that Rcmp improves the performance of RDMA-based systems via CXL, and leverages RDMA to overcome CXL’s distance limitation. To address the challenges of the mismatch between RDMA and CXL in terms of granularity, communication, and performance, Rcmp (1) provides a global page-based memory space management and enables fine-grained data access, (2) designs an efficient communication mechanism to avoid communication blocking issues, (3) proposes a hot-page identification and swapping strategy to reduce RDMA communications, and (4) designs an RDMA-optimized RPC framework to accelerate RDMA transfers. We implement a prototype of Rcmp and evaluate its performance by using micro-benchmarks and running a key-value store with YCSB benchmarks. The results show that Rcmp can achieve 5.2× lower latency and 3.8× higher throughput than RDMA-based systems. We also demonstrate that Rcmp can scale well with the increasing number of nodes without compromising performance. Zhonghua Wang 0001, Yixing Guo, Kai Lu 0002, Jiguang Wan 0001, Daohui Wang, Ting Yao 0001, Huatao Wu |
ACM Trans. Archit. Code Optim. | 5 |
| 2023 | DoW-KV: A DPU-offloaded and Write-optimized Key-Value Store on Disaggregated Persistent MemoryabstractDisaggregated Persistent Memory (DPM) is a promising technology offering elasticity, high resource utilization, persistent data storage, and lower power consumption. While building KV stores on the DPM benefits from these merits, achieving efficient writes also faces two primary challenges: 1) limited scalability caused by the underused PM bandwidth, and 2) limited CPU resources the persistent memory server (PMS) can provide. Integrating the SmartNIC such as the Data Processing Unit (DPU) into the DPM gives developers the chance to optimize writing to KV stores by utilizing both the memory and processor of DPU. However, simple offloading cannot make full use of the DPU’s potential capacity. To address these challenges, we propose DoW-KV, a persistent hash KV store on DPM. DoW-KV employs a two-tier hash index consisting of a DPU cache table in DPU memory and multiple PM persistent tables on the PM. It relocates small random writes to the DPU memory and consolidates them to the PM at a coarse granularity. Furthermore, DoW-KV uses DPU-offloaded step merge and a coroutine-based asynchronous processing framework to efficiently manage the PM persistent tables. DoW-KV also introduces a client-mixed read strategy to boost key searching on the two-tier hash index. Experimental results show that DoW-KV outperforms the state-of-the-art DINOMO by 2.1× and 1.3× in the Put and Get operations, respectively. Guokuan Li, Jiguang Wan 0001, Junyue Wang, Ting Yao 0001, Huatao Wu, Daohui Wang |
CLUSTER | 8 |
| 2023 | PetaKV: Building Efficient Key-Value Store for File System Metadata on Persistent MemoryabstractPrevious works proposed building file systems and organizing the metadata with KV stores because KV stores handle entries of various sizes efficiently and have excellent scalability. The emergence of the byte-addressable persistent memory (PM) enables metadata service to be faster than before by tailoring the KV store for the PM. However, existing PM-based KV stores cannot handle the workloads of file systems’ metadata well because simply depending on hash tables or trees cannot simultaneously provide fast file accessing and efficient directory traversing. In this paper, we exploit the insight of the metadata operations and propose the PetaKV, a KV store tailored for the metadata management of file systems on PM. PetaKV leverages dual hash indexing to achieve fast file put and get operations. Moreover, it cooperates with PM-tailored peta logs to collocate KV entries for each directory, thus supporting efficient directory scans. Our evaluation indicates PetaKV outperforms state-of-art tree-based KV stores on put, get and scan$2.5\times$,$3.2\times$, and$2.8\times$on average, respectively. Moreover, the file system built with PetaKV achieves$1.2\times$to$6.4\times$speedup compared to those built with tree-based KV stores on the metadata operations. Jian Zhou 0004, Xinhao Min, Jiguang Wan 0001, Ting Yao 0001, Daohui Wang |
IEEE Trans. Parallel Distributed Syst. | 7 |