EDBT 2026 Demo / reviewers in the wild / expert
Yudi Qiu
dblp:307/9882
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0001-6770-3681ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CXLock: Efficient and Scalable Lock Management for CXL-Enabled Distributed SystemsabstractEfficient lock management is critical in distributed systems with shared resources, especially in the context of enhanced computational power and reduced processing time. While numerous studies have leveraged Remote Direct Memory Access (RDMA) for distributed locking, these network-based approaches suffer from high network latency and jitter. The emergence of the Compute Express Link (CXL) protocol provides high-speed, low-latency connections between host processors and memory devices, presenting a compelling alternative for the lock system design. This paper introduces CXLock, the industry’s first distributed lock management system based on CXL 2.0 switches. CXLock decouples lock management into client-side queuing and manager-side arbitration, resolving race conditions of multiple clients without using remote atomics. To handle the lack of hardware-enforced coherency in CXL 2.0, CXLock employs a software-based coherency model that enables memory sharing across multiple hosts with cacheable and write-back attributes. CXLock is implemented and evaluated on a real rack-scale hardware platform containing a CXL switch and multiple host nodes. The results show that CXLock delivers an average throughput improvement of 3.7× over the RDMA-based solutions. Moreover, CXLock reduces the lock grant latency by up to 76% in both contention-free and contended scenarios. Yudi Qiu, Qixiao Liu, Wenpu Hu, Xinjun Yang, Yingqiang Zhang, Hao Chen 0080, Zipeng Ouyang, Yuemin Wu |
IEEE Trans. Computers | 2 |
| 2025 | GATe: Efficient Graph Attention Network Acceleration With Near-Memory ProcessingabstractGraph Attention Network (GAT) has gained widespread adoption thanks to its exceptional performance in processing non-Euclidean graphs. The critical components of a GAT model involve aggregation and attention, which cause numerous main-memory access, occupying significant inference time. Recently, much research has proposed near-memory processing (NMP) architectures to accelerate aggregation. However, graph attention requires additional operations distinct from aggregation, making previous NMP architectures less suitable for supporting GAT, as they typically target aggregation-only workloads. In this paper, we propose GATe, a practical and efficientGATaccelerator with NMP architecture. To the best of our knowledge, this is the first time that accelerates both attention and aggregation computation on DIMM. We unify feature vector access to eliminate the two repetitive memory accesses to source nodes caused by the sequential phase-by-phase execution of attention and aggregation. Next, we refine the computation flow to reduce data dependencies in concatenation and softmax, which lowers on-chip memory usage and communication overhead. Additionally, we introduce a novel sharding method that enhances data reusability of high-degree nodes. Experiments show that GATe achieves substantial speedup of GAT attention and aggregation phases up to 6.77× and 2.46×, with average to 3.69× and 2.24×, respectively, compared to state-of-the-art NMP works GNNear and GraNDe. Shiyan Yi, Yudi Qiu, Guohao Xu, Lingfei Lu, Xiaoyang Zeng, Yibo Fan |
IEEE Trans. Computers | 2 |
| 2025 | Flips: A Flexible Partitioning Strategy Near Memory Processing Architecture for Recommendation SystemabstractPersonalized recommendation systems are massively deployed in production data centers. The memory-intensive embedding layers of recommendation systems are the crucial performance bottleneck, with operations manifesting as sparse memory lookups and simple reduction computations. Recent studies propose near-memory processing (NMP) architectures to speed up embedding operations by utilizing high internal memory bandwidth. However, these solutions typically employ a fixed vector partitioning strategy that fail to adapt to changes in data center deployment scenarios and lack practicality. We propose Flips, aflexiblepartitioningstrategy NMP architecture that accelerates embedding layers. Flips supports more than ten partitioning strategies through hardware-software co-design. Novel hardware architectures and address mapping schemes are designed for the memory-side and host-side. We provide two approaches to determine the optimal partitioning strategy for each embedding table, enabling the architecture to accommodate changes in deployment scenarios. Importantly, Flips is decoupled from the NMP level and can utilize rank-level, bank-group-level and bank-level parallelism. In peer-level NMP evaluations, Flips outperforms state-of-the-art NMP solutions, RecNMP, TRiM, and ReCross by up to 4.0×, 4.1×, and 3.5×, respectively. Yudi Qiu, Lingfei Lu, Shiyan Yi, Minge Jing, Xiaoyang Zeng, Yibo Fan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | Gem5Tune: A Parameter Auto-Tuning Framework for Gem5 Simulator to Reduce ErrorsabstractComputer architecture simulators are widely used to explore new architectures, e.g., the gem5 simulator. However, gem5 has significant performance errors that may lead to misleading research results. Researchers typically reduce errors with the target machine by manual calibration methods, which are time-consuming and require significant expertise. This paper presents gem5Tune, a parameter auto-tuning framework for the gem5 simulator to reduce errors. Applying black-box optimization (BBO) methods, recommended for TPE-based Bayesian optimization, gem5Tune minimizes the error between gem5 and the target machine within a limited number of iterations. Three optimization methods, instruction calibration, sensitivity analysis, and dynamic pruning, are proposed to accelerate the error convergence. Experimental results show that compared to the manual calibration method, gem5Tune significantly reduces performance errors between gem5 and three modern ARM servers by more than 10% (13.83%, 10.86%, and 25.22%, respectively) for SPEC CPU benchmarks. It also scales effectively to PARSEC and SPLASH-2x benchmarks and reduces the errors of architectural events. Yudi Qiu, Xulin Yu, Xiaoyang Zeng, Yibo Fan |
IEEE Trans. Computers | 1 |
| 2024 | Scalable short-entry dual-grain coherence directories with flexible region granularity
Yudi Qiu, Jie Jiao, Yibo Fan |
J. Supercomput. | 2 |
| 2023 | Tag-Sharer-Fusion Directory: A Scalable Coherence Directory With Flexible Entry FormatsabstractIn large-scale chip multiprocessors (CMPs), the scalability of a coherence directory becomes more important as the number of cores increases. However, previously proposed scalable coherence directories typically reduce the directory storage overhead at the cost of one or more aspects of performance, accuracy, and complexity. In this article, we propose the tag-sharer-fusion (TSF) directory, a scalable coherence directory with low hardware complexity, as well as with high performance and accuracy. Each directory entry has just enough bits to store a single sharer pointer and is divided into two primary formats:tagandsharer, wheresharerentries store sharers but not tags. Each private block is tracked by atagentry, and each shared block is tracked by a combination of atagentry and asharerentry in the same set. Simulation of a 128-core chip-multiprocessor with the PARSEC and SPLASH-2x benchmarks shows that the TSF directory requires only a quarter of the area of a non-scalable full-map sparse directory to achieve similar performance and network traffic, both with an average overhead within 1%. The TSF directory outperforms the state-of-the-art Pool and way-combining directory proposals in terms of storage overhead, performance, and network traffic. Yudi Qiu, Jie Jiao, Xiaoyang Zeng, Yibo Fan |
IEEE Trans. Parallel Distributed Syst. | 1 |