VLDB 2026 Research / reviewers in the wild / expert
Wenhai Lin
dblp:257/1292
· DBLP profile ↗
14ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 2 first-author · 10 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | I-POP: Ignite Positive PrefetchersabstractHardware prefetching is a well-established technique for bridging the processor-memory performance gap. To improve cache miss coverage, modern processors often integrate multiple prefetchers. However, multi-prefetcher systems without proper management often suffer from suboptimal performance due to a surge of useless prefetches. Several techniques have been proposed to select appropriate prefetchers for issuing requests, but they all face limitations. Specifically, existing (1) static schemes lack feedback regulation mechanisms and suffer from inflexible prefetcher selections; (2) reinforcement learning (RL)based schemes incur high overhead and suffer from adjustment lag; and (3) performance-counter-based schemes rely on inefficient runtime metrics that fail to accurately and clearly reflect a prefetcher's true impact on performance. In this paper, we propose I-POP, a high-performance and lowoverhead prefetcher management scheme for multi-prefetcher systems. I-POP introduces a novel runtime metric, Prefetch Effectiveness (PE), which aggregates each prefetch request's beneficial and harmful effects to precisely quantify the impact of a prefetcher on performance, effectively overcoming the limitations of prior metrics. To compute and leverage this metric, I-POP incorporates two key components: the Metric Collector, which periodically calculates each prefetcher's PE, and the Control Engine, which dynamically manages all prefetchers based on their PE values. Specifically, I-POP ignites (enables) prefetchers with positive PE values, adaptively tuning their aggressiveness, and disables those with non-positive PE. We evaluated I-POP on numerous workloads, and the results show I-POP outperforms two state-of-the-art approaches, Bandit and Alecto, by$\mathbf{4. 2 \%}$and 3.5 % across three benchmark suites in a single-core system, and 6.6 % and 8.6 % in a 16 -core system, while incurring only 1.46 KB of storage overhead. Yiquan Lin, Wenhai Lin, Yiquan Chen, Jiexiong Xu, Shishun Cai, Jiarong Ye, Zonghui Wang, Wenzhi Chen |
HPCA | 2 |
| 2025 | OS2G: A High-Performance DPU Offloading Architecture for GPU-based Deep Learning with Object StorageabstractObject storage is increasingly attractive for deep learning (DL) applications due to its cost-effectiveness and high scalability. However, it exacerbates CPU burdens in DL clusters due to intensive object storage processing and multiple data movements. Data processing unit (DPU) offloading is a promising solution, but naively offloading the existing object storage client leads to severe performance degradation. Besides, only offloading the object storage client still involves redundant data movements, as data must first transfer from the DPU to the host and then from the host to the GPU, which continues to consume valuable host resources. Zhen Jin 0008, Yiquan Chen, Mingxu Liang, Guoju Fang, Keyao Zhang, Jiexiong Xu, Wenhai Lin, Yiquan Lin, Shushu Zhao, Wenkai Shi, Zhenhua He, Shishun Cai, Wenzhi Chen |
ASPLOS (2) | 9 |
| 2025 | rInfer: A Generic and High-Performance Framework for Remote Inference with Heterogeneous AcceleratorsabstractInference applications leverage various heterogeneous accelerators, including GPUs, TPUs, and FPGAs, to achieve remarkable performance. Remote inference, which enables applications and accelerators to reside on different nodes, can effectively address low hardware resource utilization and enhance flexibility in resource allocation. However, existing APIremoting solutions face compatibility issues, as they demand extensive development efforts to support diverse heterogeneous accelerators and adapt to rapidly evolving device runtimes. On the other hand, current inference service systems demand cumbersome cross-server configuration and suffer from suboptimal data transmission performance, overlooking data copy overhead in the network stack and not leveraging high-performance RDMA. In this paper, we present rInfer, a generic and highperformance framework for remote inference with heterogeneous accelerators. The key idea of rInfer is to abstract the entire remote inference process and optimize network transmission. On the client node, rInfer offers users an inference-oriented rInfer Device and an ease-of-use rInfer programming model, shielding users from the complexities of server configuration and network communication. On the server node, rInfer integrates with diverse inference frameworks, such as TensorRT and Torch, ensuring high compatibility with various accelerators. Furthermore, we optimize network data transmission by eliminating the need for serialization and deserialization in Protobuf to minimize data copying and adopt RDMA to enhance data transfer efficiency. Experimental results demonstrate that rInfer achieves performance nearly equivalent to local execution for the BERT model, with only a minor difference of 1.5 %. Compared to TorchServe, rInfer-TCP achieves a 51.6 % reduction in execution time for the ResNet models and consumes 37.3% less CPU and 61.2% less memory bandwidth on the server node. Zhen Jin 0008, Yiquan Chen, Yin Du, Keyao Zhang, Jiexiong Xu, Wenhai Lin, Jingchang Qin, Kanghua Fang, Wenzhi Chen |
CCGrid | 7 |
| 2025 | NVMePass: A Lightweight, High-performance and Scalable NVMe Virtualization Architecture with I/O Queues PassthroughabstractMost data-intensive applications currently run on NVMe storage, and virtualization is essential in cloud computing. Existing NVMe virtualization technologies include software-based and hardware-assisted. Virtio suffers from severe performance degradation, and polling-based solutions consume too many valuable CPU resources. Hardware-assisted solutions provide high performance and no CPU usage but have the challenges of developing dedicated hardware.In this paper, we propose NVMePass, a novel software-hardware co-design NVMe passthrough virtualization architecture designed to achieve high performance and no CPU overhead while maintaining high scalability. The key ideas of NVMePass are NVMe I/O queues passthrough for VMs and a mechanism to ensure security. The NVMePass supports DMA and interrupts remapping for VMs without hypervisor involvement, eliminating virtualization overhead and providing near-native performance. Isolation is achieved by I/O queues and logical block address resources exclusively allocated to VMs. We propose NVMe Resource Domain (NRD) and implement it in the NVMe controller to intercept illegal I/O requests. Thus, isolation and security are fully achieved. Results from our experiments show that NVMePass can provide comparable performance to VFIO, with an IOPS of $\mathbf{1 0 0. 1 \% - 1 0 0. 5 \%}$ of VFIO. Furthermore, compared to SPDK-Vhost, NVMePass achieves $\mathbf{4 0. 0 \%}$ lower latency when running 150 VMs, and NVMePass has an improvement of $\mathbf{6 8. 0 \%}$ OPS performance in a real-world application when running 100 VMs. Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Hao Yu 0016, Wenhai Lin, Kanghua Fang, Keyao Zhang, Chengkun Wei, Yuan Xie 0001, Wenzhi Chen |
HPCA | 8 |
| 2025 | Multi-Link Operation in Heterogeneous Wi-Fi 7 Networks: Modeling and Throughput OptimizationabstractMulti-link operation (MLO) is regarded as one of the most disruptive features in the upcoming IEEE 802.11be standard, known as Wi-Fi 7. However, the performance characterization of heterogeneous multi-link IEEE 802.11be networks, which consist of Multi-Link Devices (MLDs) and legacy Single-Link Devices (SLDs), remains largely unknown. The challenge originates from the lack of proper modeling of multi-link channel access schemes. In this paper, a novel model is established to study the throughput optimization of heterogeneous two-link IEEE 802.11be networks. MLDs adopt one representative synchronous multi-link channel access scheme with the primary channel. Based on the proposed model, explicit expressions of throughput of MLDs and legacy SLDs are both characterized and verified by simulation results. The network throughput is further maximized by optimally choosing the transmission probabilities of SLDs and MLDs. The analysis shows that MLO can enable MLDs to achieve higher device throughput than SLDs, yet the maximum network throughput of heterogeneous networks decreases compared to homogeneous networks composed solely of MLDs or legacy SLDs. Wenhai Lin, Xinghua Sun, Wen Zhan, Yuan Jiang 0008 |
WCNC | 1 |
| 2025 | Harmonious Coexistence Between Aloha and CSMA: Novel Dual-Channel Modeling and Throughput OptimizationabstractThe scarcity of the licensed spectrum is forcing emerging Internet of Things (IoT) networks to operate within the unlicensed spectrum. Yet there has been extensive observation indicating that performance deterioration and significant unfairness would arise, when newly deployed Aloha-based networks coexist with incumbent Carrier Sense Multiple Access (CSMA)-based WiFi networks, especially without proper adjustment of packet transmission times. The key to ensuring harmonious coexistence lies in properly modeling the coexisting networks and identifying optimal access parameters. However, the complex interactions between Aloha and CSMA nodes present significant challenges to existing analytical models. In this paper, we develop a novel dual-channel analytical framework to capture these interactions. Although Aloha and CSMA nodes coexist on the same physical channel, the developed framework represents them as operating on two separate logical channels to decouple their interactions. A discrete-time Markov renewal process is then employed to characterize the dynamics of this dual-channel network. Based on this framework, the throughput performance of the coexisting network is characterized under various packet transmission times. To achieve harmonious coexistence, the total throughput of the coexisting network under a given desired throughput proportion is optimized by tuning the packet transmission time of CSMA nodes and transmission probabilities. The optimization results indicate that the packet transmission time of CSMA nodes should be set slightly less than that of Aloha nodes. The proposed framework is further applied to enhance the network throughput and fairness of the cohabitation of LTE Unlicensed and WiFi networks. Wenhai Lin, Xinghua Sun, Anshan Yuan, Yayu Gao |
IEEE Trans. Commun. | 1 |
| 2024 | CINDA: Don't Ignore Instructions When Cloning Memory Access BehaviorabstractExisting workload cloning methods suffer from low accuracy as they primarily focus on data access patterns and ignore instruction access. This limitation reduces the accuracy of shared L2 cache design exploration and impedes processor designers from optimizing Icache and ITLB designs. In this paper, we propose CINDA, a novel workload cloning technique that can Capture both INstruction and DAta access patterns of applications. In particular, CINDA separates the instruction and data traces of applications to generate proxy instruction and proxy data traces, subsequently merging them. The results show that CINDA can accurately replicate memory access behavior with 99.1%, 99.9%, and 96.2% accuracy in replicating L1 Icache, ITLB and L2 cache performance, respectively. Furthermore, CINDA outperforms the state-of-the-art methods by reducing 7.7% L2 cache miss error. Wenhai Lin, Yiquan Chen, Jiexiong Xu, Zhen Jin 0008, Peiyu Liu 0003, Shishun Cai, Yuzhong Zhang, Jingchang Qin, Yiquan Lin, Wenzhi Chen |
CCGrid | 1 |
| 2024 | BlueJay: A Platform to Quantifying the Impact of Memory Latency on Datacenter Application PerformanceabstractUnderstanding the impact of memory latency on datacenter application performance can provide decision support to memory subsystem designers. Currently, various methods are available to quantify this impact, including cycle-accurate simulators, memory-level parallelism models, and software delay injection techniques. However, these methods suffer from several limitations, such as slow simulation speed, inaccuracy, and insufficient compatibility that requires application modification.This paper proposes BlueJay, a novel platform to quantify the impact of memory latency on the end-to-end performance of datacenter applications, avoiding slow simulation and providing high accuracy and compatibility. The key idea of BlueJay is to control the consumed memory bandwidth and read/write ratio, thereby manipulating memory latency to achieve quantification. Experiment shows that BlueJay provides accurate quantification with an average error of 3.04%. In addition, we built regression models for five applications deployed at scale in Alibaba data centers. The results reveal that a 10 ns increase in memory latency results in a performance decrease of 2.61%-3.31% for enterprise Java applications and databases, while the elastic block storage service experiences a more modest performance decrease of 0.73%-0.92%. Jingchang Qin, Yiquan Chen, Shishun Cai, Wenhai Lin, Jiexiong Xu, Zhen Jin 0008, Lifa Cao, Yuzhong Zhang, Wenzhi Chen |
CCGrid | 4 |
| 2024 | Performance Characterization of SmartNIC NVMe-over-Fabrics Target OffloadingabstractThe NVMe-over-Fabrics (NVMe-oF) is gaining popularity in cloud data centers as a remote storage protocol for accessing NVMe storage devices across servers. With the rapid increase in throughput of the NVMe storage devices, the NVMe-oF stack consumes a significant amount of valuable CPU resources. To release these valuable computing power to other tasks, many smartNICs now support NVMe-oF Target offloading. However, this emerging NVMe-oF offloading scheme's performance has not been fully investigated. Jiexiong Xu, Yiquan Chen, Wenhai Lin, Yiquan Lin, Shushu Zhao, Wenzhi Chen |
SYSTOR | 5 |
| 2024 | PARS: A Pattern-Aware Spatial Data Prefetcher Supporting Multiple Region SizesabstractHardware data prefetching is a well-studied technique to bridge the processor-memory performance gap. Bit-pattern-based prefetchers are one of the most promising spatial data prefetchers that achieve substantial performance gains. In bit-pattern-based prefetchers, the region size is a crucial parameter, which denotes the memory size that can be recorded by a pattern or prefetched by a prediction. However, existing bit-pattern-based prefetchers only support one fixed region size. Our experiment shows that the fixed region size cannot meet the requirements for numerous applications and leads to suboptimal performance and high hardware overhead. In this article, we propose PARS, a pattern-aware spatial data prefetcher supporting multiple region sizes. The key idea of PARS is that it supports multiple region sizes, enabling it to simultaneously enhance application performance while reducing the hardware overhead. Moreover, PARS supports dynamically switching appropriate region sizes for different patterns through an adaptive RS-switching mechanism. We evaluated PARS on numerous workloads and results show that PARS provides an average performance improvement of 40.6% over a baseline with no data prefetchers and outperforms the two state-of-the-art prefetchers Bingo by 2.1% (up to 24.4%) and Pythia by 3.9% (up to 111.2%) in the single-core system. In the four-core system, PARS outperforms Bingo by 5.0% (up to 66.0%) and Pythia by 5.4% (up to 177.9%). Yiquan Lin, Wenhai Lin, Jiexiong Xu, Yiquan Chen, Zhen Jin 0008, Jingchang Qin, Shishun Cai, Yuzhong Zhang, Zonghui Wang, Wenzhi Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | HyQ: Hybrid I/O Queue Architecture for NVMe over Fabrics to Enable High- Performance Hardware OffloadingabstractNVMe over Fabrics (NVMe-oF) has been widely applied as a remote storage protocol in cloud computing. The existing NVMe-oF software stack consumes a large number of CPU resources. Emerging devices, such as Smart NICs and DPUs, have supported hardware offloading of NVMe-oF to free these valuable CPU cores. However, NVMe-oF offloading capacity is always compromised because of limited hardware resources on design. Additionally, from thorough evaluations, we found that NVMe-oF inevitably suffers from severe performance degradation on complex application I/O patterns when using hardware offloading. It is challenging to achieve high performance and fully utilize NVMe-oF offloading simultaneously. In this paper, we propose HyQ, a novel hybrid I/O queue architecture for NVMe-oF, to achieve high performance while gaining the advantages of hardware offloading. HyQ realizes the coexistence of hardware offloading and software non-offloading queues, thus enabling the dynamic dispatching of I/O requests to appropriate processing queues according to user-defined I/O scheduling policies. Additionally, HyQ provides a request scheduling framework to support customized schedulers that select appropriate queues for I/O requests. In our evaluation, HyQ achieves up to 1.91x IOPS and 8.36x bandwidth performance improvement over the original hardware offloading scheme. Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Guoju Fang, Wenhai Lin, Chengkun Wei, Wenzhi Chen |
CCGrid | 8 |
| 2023 | JACO: JAva Code Layout Optimizer Enabling Continuous Optimization without Pausing Application ServicesabstractMany Java applications in data centers suffer from severe processor pipeline frontend bottlenecks, which can be mitigated by profile-guided code layout optimizations (PGCLO). To maximize optimization opportunities, state-of-the-art PGCLO solutions adopt continuous optimization to ensure that the code layout consistently matches ever-changing application control flow characteristics. However, existing continuous optimizations inevitably pause the application to execute the new code completely, which leads to high response latency and significantly deteriorates user experience.In this paper, we propose JACO, a novel profile-guided Java code layout optimizer, enabling continuous optimization without pausing application services. The key idea of JACO is to enable the execution of both the old and new code simultaneously rather than completely switching to the new code. In particular, JACO is composed of three components: (1) A lightweight profiler captures the control flow information of the application and then generates an optimized function order. (2) A control flow switcher generates new code based on optimized function order and switches the application to execute the new code without pausing the application services. (3) A selective code reclaimer only frees the memory occupied by the inactive old code. We evaluated JACO on both open-source applications and real-world applications from a world-leading company. JACO achieved up to a 16.36% performance improvement for real-world applications. The state-of-the-art approach introduces up to 37.93x latency overhead that will interrupt application services, while JACO only introduces a negligible 7% latency overhead. Wenhai Lin, Jingchang Qin, Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Yuzhong Zhang, Shishun Cai, Lirong Fu, Wenzhi Chen |
CLUSTER | 1 |
| 2022 | Deep Reinforcement Learning based Rate Adaptation for Wi-Fi NetworksabstractThe rate adaptation (RA) algorithm, which adaptively selects the rate according to the quality of the wireless environment, is one of the cornerstones of the wireless systems. In Wi-Fi networks, dynamic wireless environments are mainly due to fading channels and collisions caused by random access protocols. However, existing RA solutions mainly focus on the adaptive capability of fading channels, resulting in conservative RA policies and poor overall performance in highly congested networks. To address this problem, we propose a model-free deep reinforcement learning (DRL) based RA algorithm, named as drl RA, in this work, which incorporates the impact of collisions into the reward function design. Numerical results show that the proposed algorithm improves the throughput by 16.5% and 39.5% while reducing the latency by 25% and 19.3% compared to state-of-the-art baselines. Wenhai Lin, Peng Liu 0047, Mingjun Du, Xinghua Sun, Xun Yang 0009 |
VTC Fall | 1 |
| 2019 | Exit-Less Hypercall: Asynchronous System Calls in Virtualized Processes
Guoxi Li, Wenhai Lin, Wenzhi Chen |
ICA3PP (1) | 2 |