EDBT 2026 Demo / reviewers in the wild / expert
Yiquan Chen
dblp:241/4298
· DBLP profile ↗
17ranked-venue papers
3as first author
16since 2021 · last 2026
0009-0004-8714-6949ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 3 first-author · 14 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ada-Store: An Adaptive Load-aware Hybrid Storage Architecture for Bursty WorkloadsabstractThe adoption of quad-level cell (QLC) flash memory in NVMe solid-state drives (SSDs) offers cloud vendors high capacity and cost efficiency, but its limited performance and endurance remain critical challenges. Hybrid storage architectures (e.g., Cloud Storage Acceleration Layer (CSAL)) combine high-performance and high-density SSDs to balance performance and cost. They typically take the former as the write cache. Meanwhile, they proactively migrate data from high-performance SSDs to high-density SSDs, a process also known as compaction. However, they still face two key issues in production environments. First, they suffer from severe performance degradation under burst I/O traffic due to continuous garbage collection (GC) for high-density SSDs. Second, they adopt a fixed-size simple moving average on the compaction space reclamation rate as the upper limit of user I/O bandwidth, which can lead to user I/O performance fluctuation or poor responsiveness to rapid changes in compaction throughput. Luyang Ni, Jiexiong Xu, Yiquan Chen, Wenzhi Chen |
CF | 3 |
| 2026 | I-POP: Ignite Positive PrefetchersabstractHardware prefetching is a well-established technique for bridging the processor-memory performance gap. To improve cache miss coverage, modern processors often integrate multiple prefetchers. However, multi-prefetcher systems without proper management often suffer from suboptimal performance due to a surge of useless prefetches. Several techniques have been proposed to select appropriate prefetchers for issuing requests, but they all face limitations. Specifically, existing (1) static schemes lack feedback regulation mechanisms and suffer from inflexible prefetcher selections; (2) reinforcement learning (RL)based schemes incur high overhead and suffer from adjustment lag; and (3) performance-counter-based schemes rely on inefficient runtime metrics that fail to accurately and clearly reflect a prefetcher's true impact on performance. In this paper, we propose I-POP, a high-performance and lowoverhead prefetcher management scheme for multi-prefetcher systems. I-POP introduces a novel runtime metric, Prefetch Effectiveness (PE), which aggregates each prefetch request's beneficial and harmful effects to precisely quantify the impact of a prefetcher on performance, effectively overcoming the limitations of prior metrics. To compute and leverage this metric, I-POP incorporates two key components: the Metric Collector, which periodically calculates each prefetcher's PE, and the Control Engine, which dynamically manages all prefetchers based on their PE values. Specifically, I-POP ignites (enables) prefetchers with positive PE values, adaptively tuning their aggressiveness, and disables those with non-positive PE. We evaluated I-POP on numerous workloads, and the results show I-POP outperforms two state-of-the-art approaches, Bandit and Alecto, by$\mathbf{4. 2 \%}$and 3.5 % across three benchmark suites in a single-core system, and 6.6 % and 8.6 % in a 16 -core system, while incurring only 1.46 KB of storage overhead. Yiquan Lin, Wenhai Lin, Yiquan Chen, Jiexiong Xu, Shishun Cai, Jiarong Ye, Zonghui Wang, Wenzhi Chen |
HPCA | 3 |
| 2026 | Spillway: Orchestrating DPU and Host into a Unified vSwitching FabricabstractThe transition to Data Processing Unit (DPU)-centric architectures has become the de-facto standard in modern cloud networks, enabling infrastructure offload and improved host resource utilization. However, the fixed hardware limits of DPUs increasingly fail to keep pace with the rapid growth of host compute density and network-intensive workloads. As a result, when DPU resources are saturated, host compute capacity often remains underutilized due to insufficient network provisioning. Xiaochong Jiang, Yilong Lv, Naixuan Guan, Qiming Zhao, Sihan Fu, Xuyang Ge, Denghui Wu, Yibin Shen, Guochun Hong, Yijian Dong, Yiquan Chen, Shaoliang An, Zhixiong Guo, Yisong Qiao, Hongwei Ding 0004, Shize Zhang, Rong Wen, Yang Song 0031, Zhigang Zong, Xing Li 0007, Chengkun Wei, Shunmin Zhu, Wenzhi Chen |
SIGCOMM | 15 |
| 2025 | OS2G: A High-Performance DPU Offloading Architecture for GPU-based Deep Learning with Object StorageabstractObject storage is increasingly attractive for deep learning (DL) applications due to its cost-effectiveness and high scalability. However, it exacerbates CPU burdens in DL clusters due to intensive object storage processing and multiple data movements. Data processing unit (DPU) offloading is a promising solution, but naively offloading the existing object storage client leads to severe performance degradation. Besides, only offloading the object storage client still involves redundant data movements, as data must first transfer from the DPU to the host and then from the host to the GPU, which continues to consume valuable host resources. Zhen Jin 0008, Yiquan Chen, Mingxu Liang, Guoju Fang, Keyao Zhang, Jiexiong Xu, Wenhai Lin, Yiquan Lin, Shushu Zhao, Wenkai Shi, Zhenhua He, Shishun Cai, Wenzhi Chen |
ASPLOS (2) | 2 |
| 2025 | rInfer: A Generic and High-Performance Framework for Remote Inference with Heterogeneous AcceleratorsabstractInference applications leverage various heterogeneous accelerators, including GPUs, TPUs, and FPGAs, to achieve remarkable performance. Remote inference, which enables applications and accelerators to reside on different nodes, can effectively address low hardware resource utilization and enhance flexibility in resource allocation. However, existing APIremoting solutions face compatibility issues, as they demand extensive development efforts to support diverse heterogeneous accelerators and adapt to rapidly evolving device runtimes. On the other hand, current inference service systems demand cumbersome cross-server configuration and suffer from suboptimal data transmission performance, overlooking data copy overhead in the network stack and not leveraging high-performance RDMA. In this paper, we present rInfer, a generic and highperformance framework for remote inference with heterogeneous accelerators. The key idea of rInfer is to abstract the entire remote inference process and optimize network transmission. On the client node, rInfer offers users an inference-oriented rInfer Device and an ease-of-use rInfer programming model, shielding users from the complexities of server configuration and network communication. On the server node, rInfer integrates with diverse inference frameworks, such as TensorRT and Torch, ensuring high compatibility with various accelerators. Furthermore, we optimize network data transmission by eliminating the need for serialization and deserialization in Protobuf to minimize data copying and adopt RDMA to enhance data transfer efficiency. Experimental results demonstrate that rInfer achieves performance nearly equivalent to local execution for the BERT model, with only a minor difference of 1.5 %. Compared to TorchServe, rInfer-TCP achieves a 51.6 % reduction in execution time for the ResNet models and consumes 37.3% less CPU and 61.2% less memory bandwidth on the server node. Zhen Jin 0008, Yiquan Chen, Yin Du, Keyao Zhang, Jiexiong Xu, Wenhai Lin, Jingchang Qin, Kanghua Fang, Wenzhi Chen |
CCGrid | 2 |
| 2025 | NVMePass: A Lightweight, High-performance and Scalable NVMe Virtualization Architecture with I/O Queues PassthroughabstractMost data-intensive applications currently run on NVMe storage, and virtualization is essential in cloud computing. Existing NVMe virtualization technologies include software-based and hardware-assisted. Virtio suffers from severe performance degradation, and polling-based solutions consume too many valuable CPU resources. Hardware-assisted solutions provide high performance and no CPU usage but have the challenges of developing dedicated hardware.In this paper, we propose NVMePass, a novel software-hardware co-design NVMe passthrough virtualization architecture designed to achieve high performance and no CPU overhead while maintaining high scalability. The key ideas of NVMePass are NVMe I/O queues passthrough for VMs and a mechanism to ensure security. The NVMePass supports DMA and interrupts remapping for VMs without hypervisor involvement, eliminating virtualization overhead and providing near-native performance. Isolation is achieved by I/O queues and logical block address resources exclusively allocated to VMs. We propose NVMe Resource Domain (NRD) and implement it in the NVMe controller to intercept illegal I/O requests. Thus, isolation and security are fully achieved. Results from our experiments show that NVMePass can provide comparable performance to VFIO, with an IOPS of $\mathbf{1 0 0. 1 \% - 1 0 0. 5 \%}$ of VFIO. Furthermore, compared to SPDK-Vhost, NVMePass achieves $\mathbf{4 0. 0 \%}$ lower latency when running 150 VMs, and NVMePass has an improvement of $\mathbf{6 8. 0 \%}$ OPS performance in a real-world application when running 100 VMs. Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Hao Yu 0016, Wenhai Lin, Kanghua Fang, Keyao Zhang, Chengkun Wei, Yuan Xie 0001, Wenzhi Chen |
HPCA | 1 |
| 2025 | Alibaba Stellar: A New Generation RDMA Network for Cloud AIabstractThe rapid adoption of Large Language Models (LLMs) in cloud environments has intensified the demand for high-performance AI training and inference, where Remote Direct Memory Access (RDMA) plays a critical role. However, existing RDMA virtualization solutions, such as Single-Root Input/Output Virtualization (SR-IOV), face significant limitations in scalability, performance, and stability. These issues include lengthy container initialization times, hardware resource constraints, and inefficient traffic steering. To address these challenges, we propose Stellar, a new generation RDMA network for cloud AI. Stellar introduces three key innovations: Para-Virtualized Direct Memory Access (PVDMA) for on-demand memory pinning, extended Memory Translation Table (eMTT) for optimized GPU Direct RDMA (GDR) performance, and RDMA Packet Spray for efficient multi-path utilization. Deployed in our large-scale AI clusters, Stellar spins up virtual devices in seconds, reduces container initialization time by 15 times, and improves LLM training speed by up to 14%. Our evaluations demonstrate that Stellar significantly outperforms existing solutions, offering a scalable, stable, and high-performance RDMA network for cloud AI. Menglei Zheng, Binbin Liao, Suwei Xu, Yongjia Mo, Qinghua Peng, Jilie Luo, Qingxu Li, Zishu Wang, Jianbo Dong, Kunling He, Sheng Cheng 0002, Jiamin Cao, Hairong Jiao, Lingjun Zhu, Yiquan Chen, Wei Wang 0030, Shuhong Zhu, Xingru Li, Qiang Wang 0022, Wei Lin 0016, Ennan Zhai, Jiesheng Wu, Qiang Liu 0036, Binzhang Fu, Dennis Cai |
SIGCOMM | 27 |
| 2024 | CINDA: Don't Ignore Instructions When Cloning Memory Access BehaviorabstractExisting workload cloning methods suffer from low accuracy as they primarily focus on data access patterns and ignore instruction access. This limitation reduces the accuracy of shared L2 cache design exploration and impedes processor designers from optimizing Icache and ITLB designs. In this paper, we propose CINDA, a novel workload cloning technique that can Capture both INstruction and DAta access patterns of applications. In particular, CINDA separates the instruction and data traces of applications to generate proxy instruction and proxy data traces, subsequently merging them. The results show that CINDA can accurately replicate memory access behavior with 99.1%, 99.9%, and 96.2% accuracy in replicating L1 Icache, ITLB and L2 cache performance, respectively. Furthermore, CINDA outperforms the state-of-the-art methods by reducing 7.7% L2 cache miss error. Wenhai Lin, Yiquan Chen, Jiexiong Xu, Zhen Jin 0008, Peiyu Liu 0003, Shishun Cai, Yuzhong Zhang, Jingchang Qin, Yiquan Lin, Wenzhi Chen |
CCGrid | 2 |
| 2024 | BlueJay: A Platform to Quantifying the Impact of Memory Latency on Datacenter Application PerformanceabstractUnderstanding the impact of memory latency on datacenter application performance can provide decision support to memory subsystem designers. Currently, various methods are available to quantify this impact, including cycle-accurate simulators, memory-level parallelism models, and software delay injection techniques. However, these methods suffer from several limitations, such as slow simulation speed, inaccuracy, and insufficient compatibility that requires application modification.This paper proposes BlueJay, a novel platform to quantify the impact of memory latency on the end-to-end performance of datacenter applications, avoiding slow simulation and providing high accuracy and compatibility. The key idea of BlueJay is to control the consumed memory bandwidth and read/write ratio, thereby manipulating memory latency to achieve quantification. Experiment shows that BlueJay provides accurate quantification with an average error of 3.04%. In addition, we built regression models for five applications deployed at scale in Alibaba data centers. The results reveal that a 10 ns increase in memory latency results in a performance decrease of 2.61%-3.31% for enterprise Java applications and databases, while the elastic block storage service experiences a more modest performance decrease of 0.73%-0.92%. Jingchang Qin, Yiquan Chen, Shishun Cai, Wenhai Lin, Jiexiong Xu, Zhen Jin 0008, Lifa Cao, Yuzhong Zhang, Wenzhi Chen |
CCGrid | 2 |
| 2024 | LightPool: A NVMe-oF-based High-performance and Lightweight Storage Pool Architecture for Cloud-Native Distributed DatabaseabstractEmerging cloud-native distributed databases rely on local NVMe SSDs to provide high-performance and highavailable data services to many cloud applications. However, the database clusters suffer from low utilization of local storage because of the imbalance between CPU and storage capacities within each node. For instance, the OceanBase distributed database cluster, with hundreds of PB local storage capacity, only utilizes around 40% of its local storage. Although disaggregated storage (EBS) can enhance storage utilization by provisioning the CPU and storage independently on demand, they suffer from performance bottlenecks and high costs. In this paper, we propose LightPool, a high-performance and lightweight storage pool architecture large-scale deployed in the OceanBase clusters, enhancing storage resource utilization. The key idea of LightPool is aggregating cluster storage into a storage pool and enabling unified management. In particular, LightPool adopts NVMe-oF to enable high-performance storage resource sharing among cluster nodes and integrate the storage pool with Kubernetes to achieve flexible management and allocation of storage resources. Furthermore, we design the hot-upgrade and hot-migration mechanisms to enhance the availability of LightPool. We have deployed LightPool on over 8500 nodes in production clusters. Statistics show that LightPool can improve storage resource utilization from about 40% to 65%. Experimental results show that the extra latency from LightPool is only about 2.1 μs compared to local storage. Compared to OpenEBS, LightPool enhances bandwidth up to 190.9% in microbenchmarks and throughput up to 6.9% in real-world applications. LightPool is the best practice to deploy NVMe-oF (NVMe/TCP) in the production environment. We also discuss important lessons and experiences learned from the development of LightPool. Jiexiong Xu, Yiquan Chen, Wenhui Shi, Guoju Fang, Huasheng Liao, Zhen Jin 0008, Wenzhi Chen |
HPCA | 2 |
| 2024 | Performance Characterization of SmartNIC NVMe-over-Fabrics Target OffloadingabstractThe NVMe-over-Fabrics (NVMe-oF) is gaining popularity in cloud data centers as a remote storage protocol for accessing NVMe storage devices across servers. With the rapid increase in throughput of the NVMe storage devices, the NVMe-oF stack consumes a significant amount of valuable CPU resources. To release these valuable computing power to other tasks, many smartNICs now support NVMe-oF Target offloading. However, this emerging NVMe-oF offloading scheme's performance has not been fully investigated. Jiexiong Xu, Yiquan Chen, Wenhai Lin, Yiquan Lin, Shushu Zhao, Wenzhi Chen |
SYSTOR | 3 |
| 2024 | PARS: A Pattern-Aware Spatial Data Prefetcher Supporting Multiple Region SizesabstractHardware data prefetching is a well-studied technique to bridge the processor-memory performance gap. Bit-pattern-based prefetchers are one of the most promising spatial data prefetchers that achieve substantial performance gains. In bit-pattern-based prefetchers, the region size is a crucial parameter, which denotes the memory size that can be recorded by a pattern or prefetched by a prediction. However, existing bit-pattern-based prefetchers only support one fixed region size. Our experiment shows that the fixed region size cannot meet the requirements for numerous applications and leads to suboptimal performance and high hardware overhead. In this article, we propose PARS, a pattern-aware spatial data prefetcher supporting multiple region sizes. The key idea of PARS is that it supports multiple region sizes, enabling it to simultaneously enhance application performance while reducing the hardware overhead. Moreover, PARS supports dynamically switching appropriate region sizes for different patterns through an adaptive RS-switching mechanism. We evaluated PARS on numerous workloads and results show that PARS provides an average performance improvement of 40.6% over a baseline with no data prefetchers and outperforms the two state-of-the-art prefetchers Bingo by 2.1% (up to 24.4%) and Pythia by 3.9% (up to 111.2%) in the single-core system. In the four-core system, PARS outperforms Bingo by 5.0% (up to 66.0%) and Pythia by 5.4% (up to 177.9%). Yiquan Lin, Wenhai Lin, Jiexiong Xu, Yiquan Chen, Zhen Jin 0008, Jingchang Qin, Shishun Cai, Yuzhong Zhang, Zonghui Wang, Wenzhi Chen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | HyQ: Hybrid I/O Queue Architecture for NVMe over Fabrics to Enable High- Performance Hardware OffloadingabstractNVMe over Fabrics (NVMe-oF) has been widely applied as a remote storage protocol in cloud computing. The existing NVMe-oF software stack consumes a large number of CPU resources. Emerging devices, such as Smart NICs and DPUs, have supported hardware offloading of NVMe-oF to free these valuable CPU cores. However, NVMe-oF offloading capacity is always compromised because of limited hardware resources on design. Additionally, from thorough evaluations, we found that NVMe-oF inevitably suffers from severe performance degradation on complex application I/O patterns when using hardware offloading. It is challenging to achieve high performance and fully utilize NVMe-oF offloading simultaneously. In this paper, we propose HyQ, a novel hybrid I/O queue architecture for NVMe-oF, to achieve high performance while gaining the advantages of hardware offloading. HyQ realizes the coexistence of hardware offloading and software non-offloading queues, thus enabling the dynamic dispatching of I/O requests to appropriate processing queues according to user-defined I/O scheduling policies. Additionally, HyQ provides a request scheduling framework to support customized schedulers that select appropriate queues for I/O requests. In our evaluation, HyQ achieves up to 1.91x IOPS and 8.36x bandwidth performance improvement over the original hardware offloading scheme. Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Guoju Fang, Wenhai Lin, Chengkun Wei, Wenzhi Chen |
CCGrid | 1 |
| 2023 | JACO: JAva Code Layout Optimizer Enabling Continuous Optimization without Pausing Application ServicesabstractMany Java applications in data centers suffer from severe processor pipeline frontend bottlenecks, which can be mitigated by profile-guided code layout optimizations (PGCLO). To maximize optimization opportunities, state-of-the-art PGCLO solutions adopt continuous optimization to ensure that the code layout consistently matches ever-changing application control flow characteristics. However, existing continuous optimizations inevitably pause the application to execute the new code completely, which leads to high response latency and significantly deteriorates user experience.In this paper, we propose JACO, a novel profile-guided Java code layout optimizer, enabling continuous optimization without pausing application services. The key idea of JACO is to enable the execution of both the old and new code simultaneously rather than completely switching to the new code. In particular, JACO is composed of three components: (1) A lightweight profiler captures the control flow information of the application and then generates an optimized function order. (2) A control flow switcher generates new code based on optimized function order and switches the application to execute the new code without pausing the application services. (3) A selective code reclaimer only frees the memory occupied by the inactive old code. We evaluated JACO on both open-source applications and real-world applications from a world-leading company. JACO achieved up to a 16.36% performance improvement for real-world applications. The state-of-the-art approach introduces up to 37.93x latency overhead that will interrupt application services, while JACO only introduces a negligible 7% latency overhead. Wenhai Lin, Jingchang Qin, Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Yuzhong Zhang, Shishun Cai, Lirong Fu, Wenzhi Chen |
CLUSTER | 3 |
| 2023 | BM-Store: A Transparent and High-performance Local Storage Architecture for Bare-metal Clouds Enabling Large-scale DeploymentabstractBare-metal instances are crucial for high-value, mission-critical applications on the cloud. Tenants exclusively use these dedicated hardware resources. Local virtualized disks are essential for bare-metal instances to provide flexible and high-performance storage resources. Traditionally tenants can choose polling-based software virtualization techniques, but they consume too many valuable host CPU cores and suffer from performance degradation. Cloud vendors are hard to deploy existing hardware-assisted local storage solutions in bare-metal instances due to no access to the host OS to install customized drivers. Moreover, cloud vendors have difficulties managing and maintaining the local storage devices in bare-metal instances because hardware resources and host operating systems are completely utilized by tenants, then it will impact the availability of storage devices.This paper presents our design and experience with BM-Store, a novel high-performance hardware-assisted virtual local storage architecture for bare-metal clouds. BM-Store is transparent to the host that tenants are unaware of the underlying hardware architecture. Therefore, it can be deployed on a large scale in cloud vendors. BM-Store consists of two components: an FPGA-based BMS-Engine and an ARM-based BMS-Controller. The BMS-Engine accelerates the I/O path to enable high-performance virtual storage independent of disk devices without consuming any CPU resource on the host. The BMS-Controller is responsible for resource management and maintenance to achieve flexible and high available local storage. The results of the extensive experiments show that BM-Store can achieve near-native performance, which only introduces about 3 µs extra latency and average 4.0% throughput overhead to native disks. Compared to SPDK vhost, BM-Store achieves an average bandwidth improvement of 15.7% in microbenchmark and a maximum throughput enhancement of 13.4% in real-world applications. Yiquan Chen, Jiexiong Xu, Chengkun Wei, Xulin Yu, Zeke Wang, Shuibing He, Wenzhi Chen |
HPCA | 1 |
| 2021 | On Workload-Aware DRAM Failure Prediction in Large-Scale Data CentersabstractDRAM failures are one of the major hardware threats to the reliability of large-scale data centers since the uncorrectable errors in DRAMs may cause servers to shut down. Existing works try to solve this problem by predicting DRAM failures in advance with Machine Learning models. In these works, correctable errors (CEs) are generally deemed as the most important feature. The major reason behind CEs' emergence is the accumulated stress caused by intensive workloads. Moreover, defective DRAMs will not manifest themselves as system errors until the defective cells are accessed by some specific workloads. Therefore, the running workloads on a server are also important for DRAM failure prediction. In this paper, we focus on the impact of workloads on DRAM failures. We design the workload features from both macroscopical and microscopical aspects, i.e. node-level performance metrics and cell-level DRAM access pattern, respectively. Furthermore, we propose Hierarchical DRAM Error Code (HiDEC) to represent the DRAM access pattern. We leverage several Decision Tree-based models for DRAM failure prediction to highlight the generality of our designed features. Experiments are carried out based on the dataset collected from a real-world commercial data center. The results show that both macroscopic and microscopic features can bring significant improvements to the prediction performance. Xingyi Wang, Yu Li 0007, Yiquan Chen, Yin Du, Yuzhong Zhang, Pinan Chen, Wenjun Song, Qiang Xu 0001, Li Jiang 0002 |
VTS | 3 |
| 2019 | System-level hardware failure prediction using deep learningabstractDisk and memory faults are the leading causes of server breakdown. A proactive solution is to predict such hardware failure at the runtime and then isolate the hardware at risk and backup the data. However, the current model-based predictors are incapable of using the discrete time-series data, such as the values of device attributes, which conveys high-level information of the device behavior. In this paper, we propose a novel deep-learning based prediction scheme for system-level hardware failure prediction. We normalize the distribution of samples' attributes from different vendors to make use of diverse training sets. We propose a temporal Convolution Neural Network based model that is insensitive to the noise in the time dimension. Finally, we design a loss function to train the model with extremely imbalanced samples effectively. Experimental results from an open S.M.A.R.T data set and an industrial data set show the effectiveness of the proposed scheme. Xiaoyi Sun, Krishnendu Chakrabarty, Ruirui Huang, Yiquan Chen, Hai Cao, Yinhe Han 0001, Xiaoyao Liang, Li Jiang 0002 |
DAC | 4 |