VLDB 2026 Research / reviewers in the wild / expert
Yuke Li 0003
dblp:128/0927-3
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0009-5806-7228ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When RDMA Goes Long-Haul: Characterization, Modeling, and Verbs-Level Emulation with Implications for Federated LearningabstractLong-haul Remote Direct Memory Access (RDMA) is rapidly emerging as a viable mechanism for extending high-performance communication beyond traditional datacenter environments into wide-area networks (WANs). However, increased round-trip time (RTT) fundamentally alters RDMA behavior, shifting execution from open-loop injection toward a closed-loop, window-limited regime governed by resource constraints and concurrency effects. Consequently, performance characteristics diverge from established datacenter assumptions, rendering existing models insufficient for predicting performance at WAN scale. Understanding these dynamics requires real-world characterization, yet geographically distributed long-haul testbeds remain expensive, scarce, and difficult to reproduce. Yuke Li 0003, Zhonghao Chen, Xiaoyi Lu 0001 |
HPDC | 1 |
| 2026 | PACER: A Userspace Network Rate Controller in MPI with Adaptive Compression for Parallel Applications
Yuke Li 0003, Darren Ng, Arjun Kashyap, Sheng Di, Guanpeng Li, Xiaoyi Lu 0001 |
ICS | 1 |
| 2025 | DPU-KV: On the Benefits of DPU Offloading for In-Memory Key-Value Stores at the EdgeabstractIn-memory key-value stores (KVS) are widely used for edge data storage, where low latency and high throughput are essential. Data Processing Units (DPUs), with their low power use and offloading capabilities, suit resource-constrained edge computing. While DPUs offer a new design point for KVS, their offloading in edge environments remains underexplored and challenging. In this paper, we unveil the potential of offloading in-memory CPU-based KVS to SoC-based DPUs, specifically NVIDIA's BlueField-2 (BF-2) and BlueField-3 (BF-3), with the aim of enhancing KVS performance. We propose a principled exploration methodology of dividing a KVS (i.e., MICA) into its logical components and identifying the CPU-intensive KVS component (i.e., communication engine). Next, we perform fine-grained offloading analysis and explorations on DPUs. To maximize benefits in terms of latency and throughput from fine-grained KVS offloading on DPUs, we propose a series of significant performance optimizations, including a key-value-based queue-pair model, overlapped KV request/response processing, reduced DMA operations per KV batch, dual-communication engine, and a sharding-based design. Our key finding is that our proposed fine-grained KVS offloading designs on modern DPU architectures (i.e., BF-2 and BF-3) can provide much lower latency (up to 68%) and higher throughput (up to 36%) than MICA (CPU-only) and coarse-grained DPU offloading schemes at the edge. To our knowledge, this paper is the first to explore the performance benefits of fine-grained KVS offloading to DPUs at the edge. Arjun Kashyap, Yuke Li 0003, Xiaoyi Lu 0001 |
HPDC | 2 |
| 2025 | Understanding the Idiosyncrasies of Emerging BlueField DPUsabstractData Processing Units (DPUs) are becoming available in datacenter environments to offload/accelerate workloads from the host.However, a comprehensive analysis is required to help users determine how to effectively utilize DPUs for their workloads, considering the various configurations and generations available.To fill in this gap, we conduct a fair and rigorous characterization by performing 15 benchmarking tests to demonstrate the evolution of representative SoC-based DPUs, specifically NVIDIA's BlueField-1, BlueField-2, and BlueField-3.Our work surfaces several idiosyncrasies across three key characterization dimensions-network, DMA engine, and memory.For network, we exhaustively test two major DPU modes-on-path (and five submodes) and offpath modes.We develop DPUDMABench, a microbenchmark suite to systematically analyze different data exchange primitives supported by DPU's DMA engine.We also conduct two application case studies examining the DPU mode's performance impact on TCP/IP and RDMA-based key-value stores (MICA and HERD).Based on our multi-generational DPU characterization, we identify and summarize 14 major idiosyncrasies, along with providing guidelines for optimal system and future hardware design. Arjun Kashyap, Yuke Li 0003, Darren Ng, Xiaoyi Lu 0001 |
ICS | 2 |
| 2025 | Heliostat: Harnessing Ray Tracing Accelerators for Page Table WalksabstractThis paper introduces Heliostat, which enhances page translation bandwidth on GPUs by harnessing underutilized ray tracing accelerators (RTAs).While most existing studies focused on better utilizing the provided translation bandwidth, this paper introduces a new opportunity to fundamentally increase the translation bandwidth.Instead of overprovisioning the GPU memory management unit (GMMU), Heliostat repurposes the existing RTAs by leveraging the operational similarities between ray tracing and page table walks.Unlike earlier studies that utilized RTAs for certain workloads, Heliostat democratizes RTA for supporting any workloads by improving virtual memory performance.Heliostat+ optimizes Heliostat by handling predicted future address translations proactively.Heliostat outperforms baseline and two state-of-the-arts by 1.93×, 1.92×, and 1.66×.Heliostat+ further speeds up Heliostat by 1.23×.Compared to an overprovisioned comparable solution, Heliostat occupies only 1.53% of the area and consumes 5.8% of the power. Yuke Li 0003, Jiwon Lee 0001, Won Woo Ro, Hyeran Jeon |
ISCA | 2 |
| 2024 | Accelerating Lossy and Lossless Compression on Emerging BlueField DPU ArchitecturesabstractData compression has become a crucial technique in addressing performance bottlenecks caused by increasing data volumes in High-Performance Computing (HPC), Big Data, and Deep Learning (DL). Despite its potential to boost system performance, recent studies have identified significant challenges with existing compression methods, mainly due to their high computational demands amidst continuously growing data sizes. Concurrently, the advent of Data Processing Units (DPUs), equipped with programmable System-on-Chip (SoC) and specialized compression accelerators, offers a promising opportunity to alter the landscape of data compression. This paper explores the complexities and potential of leveraging NVIDIA BlueField DPUs to accelerate lossy and lossless compression. Towards this, we introduce PEDAL, an innovative library that leverages the hardware capabilities of DPUs to unify and optimize data compression designs. Moreover, we seamlessly co-design PEDAL with the popular MPICH MPI library, demonstrating up to 101x speedup in compression time and 88x decrease in communication latency. Drawing on these achievements, we share our experience with various research communities about accelerating data compression on DPUs in communication-oriented HPC scenarios. Yuke Li 0003, Arjun Kashyap, Weicong Chen 0002, Yanfei Guo, Xiaoyi Lu 0001 |
IPDPS | 1 |
| 2024 | On the Feasibility and Benefits of Extensive EvaluationabstractBenchmark and system parameters often have a significant impact on performance evaluation, which raises a long-lasting question about which settings we should use. This paper studies the feasibility and benefits of extensive evaluation. A full extensive evaluation, which tests all possible settings, is usually too expensive. This work investigates whether it is possible to sample a subset of the settings and, upon them, generate observations that match those from a full extensive evaluation. Towards this goal, we have explored the incremental sampling approach, which starts by measuring a small subset of random settings, builds a prediction model on these samples using the popular ANOVA approach, adds more samples if the model is not accurate enough, and terminates otherwise. To summarize our findings: 1) Enhancing a research prototype to support extensive evaluation mostly involves changing hard-coded configurations, which does not take much effort. 2) Some systems are highly predictable, which means that they can achieve accurate predictions with a low sampling rate, but some systems are less predictable. 3) We have not found a method that can consistently outperform random sampling + ANOVA. Based on these findings, we provide recommendations to improve artifact predictability and strategies for selecting parameter values during evaluation. Yujie Hui, Miao Yu 0023, Hao Qi 0008, Yifan Gan, Tianxi Li, Yuke Li 0003, Xueyuan Ren, Sixiang Ma, Xiaoyi Lu 0001, Yang Wang 0009 |
Proc. ACM Manag. Data | 6 |
| 2023 | Characterizing Lossy and Lossless Compression on Emerging BlueField DPU ArchitecturesabstractThe Data Processing Unit (DPU) (i.e., programmable SmartNICs with System-on-Chip or SoC cores) has emerged as a valuable supplementary resource to the host CPU. The DPU architecture has been attracting significant attention within High-Performance Computing (HPC) and data center clusters due to its advanced capabilities and accelerators, which include a hardware-based data compression engine. This positions the DPU as a prospective tool for accelerating and offloading compression workloads from the hosts, which can potentially speed up data-intensive applications. The convergence of Big Data, HPC, and Machine Learning (ML) systems has rendered large data volumes a major performance bottleneck in message communication and data storage. While compression can boost performance, recent studies reveal that compression techniques (e.g., lossy and lossless) are compute-intensive and time-consuming, particularly with larger data sizes. Consequently, this paper characterizes the performance of three lossy (SZ3) and lossless (DEFLATE and zlib) compression algorithms with seven real-world data sets on the popular NVIDIA’s BlueField DPUs to explore potential opportunities for offloading these workloads from the host. We find that compared to DPU’s SoC cores, DPU’s hardware compression engine can obtain up to 26.8x performance speedup. Furthermore, we discuss the challenges and opportunities associated with employing NVIDIA’s BlueField DPUs to accelerate lossy and lossless compression/decompression workloads. Our research discloses five important takeaways which shed light on future research directions for lossy and lossless compressions on DPUs. Yuke Li 0003, Arjun Kashyap, Yanfei Guo, Xiaoyi Lu 0001 |
HOTI | 1 |
| 2023 | xCCL: A Survey of Industry-Led Collective Communication Libraries for Deep Learning
Adam Weingram, Yuke Li 0003, Hao Qi 0008, Darren Ng, Liuyao Dai, Xiaoyi Lu 0001 |
J. Comput. Sci. Technol. | 2 |