Bo Peng 0043

dblp:03/5954-43 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-6190-5327ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 7 first-author · 9 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 REPA: Reconfigurable PIM for the Joint Acceleration of KV Cache Offloading and Processing
abstract
The use of KV cache in LLM inference leads to large memory footprint and sub-optimal decoding performance. Prior studies typically address one of these two limitations by either offloading or stage-split inference. In this paper, we explore and reveal the possibility of a joint solution, and propose REPA, a GPU-PIM hybrid system to prototype this idea. We leverage reconfigurable ReRAM PIM to achieve fast KV cache persistence, and balance the requirement of processing speed and memory capacity. To fully unleash the parallelization potential of REPA, we propose optimizations in (1) architecture, (2) data mapping and (3) pipelining: (1) We propose bulk-wise memory instructions and multi-level controllers to enable finer-grained parallelism in the PIM device. (2) We propose locality-aware data mapping to make the best of the aforementioned architectural optimization, and reduce long-range data transfer on chip. (3) We adopt sub-batch pipelining to reduce idleness in batches, and propose transfer overlapping to shadow the KV cache transfer by computation. Experimental results show that REPA exhibits high inference speed, energy efficiency and integratability. It is 1.5--6.5× faster, and 8--10× more efficient than NVIDIA A100. It also outperforms state-of-the-art DRAM PIM systems by up to 1.4× for long context inference. When integrated into existing offloading systems, REPA achieves 1.4--2.0× offloading speed, and 1.2--1.4× end-to-end speedup, showcasing its high potential for fast KV cache offloading and processing.
Junlong Yang, Bo Peng 0043, Jianguo Yao 0002
ASPLOS (2)3
2026 Mitigating Cold Starts in Container-Ephemeral Large Language Model Serving
Bo Peng 0043, Jianguo Yao 0002
IWQoS2
2026 Zero2M: Optimizing Tenant-Level I/O Management for Future Faster NVMe Storage with FPGA
abstract
High-speed Non-Volatile Memory Express (NVMe) Solid-State Drives (SSDs) are shared by multiple tenants in cloud scenarios to improve resource utilization. Tenant-level I/O management is necessary to achieve reliable QoS control during sharing. Unfortunately, our investigation finds that CPUs inevitably participate in I/O management for existing solutions because SSDs are not tenant-sensitive and have limited internal computing resources. It introduces additional CPU costs and latency overhead when serving future faster SSDs. We propose that the Field Programmable Logic Gate Array (FPGA) is a promising alternative for freeing tenant-level I/O management from CPUs. However, implementing tenant-level I/O management using the FPGA requires addressing the following challenges: (1) System compatibility and tenant identification; (2) Efficient FPGA workflows that will not become a bottleneck; (3) Fast I/O management workflow that introduces the lowest additional CPU costs and latency. This article presents Zero2M, a novel CPU-free system designed to optimize the additional CPU costs and latency overhead in tenant-level I/O management for future faster NVMe SSDs. Zero2M proposes a dedicated FPGA-based NVMe controller to preserve system compatibility and identify tenants using the namespace mechanism in NVMe. It allows I/O management without modifying host software, which existing solutions cannot achieve. The parallelized and pipelined workflows are proposed in the controller to accelerate I/O command processing and prevent the controller from becoming a bottleneck for the I/O management workflow. The read/write speed of the Zero2M controller is 4.65 \(\times\) /4.92 \(\times\) faster than the state-of-the-art hardware-accelerated controller. The I/O management workflow is formulated as a novel parallelized and pipelined accelerator and integrated into the workflow of Zero2M’s controller. It optimizes additional CPU costs and latency overhead for tenant-level I/O management. Experiments present that Zero2M reduces an average of 3.01 \(\times\) CPU usage while maintaining the lowest latency overhead (7.62 \(\times\) lower on average) compared to the state-of-the-art solution. It also removes the CPU dependency for tenant-level I/O management for the first time.
Bo Peng 0043, Jianguo Yao 0002, Haibing Guan
ACM Trans. Reconfigurable Technol. Syst.2
2025 Leopard: Hardware Pass-Through Remote Storage Access with Queue Concurrency for Edge Intelligent Workstations
abstract
Edge intelligent workstations load (store) massive empirical data from (to) remote cloud storage due to limited local storage. However, current remote storage access frameworks are complex. They use expensive computing resources to manipulate multiple concurrent request queues in modern high-speed storage devices for saturating performance. Complex software stacks and limited CPUs on edge intelligent workstations hinder saturating concurrent request queues, thus resulting in up to $75 \%$ performance degradation for remote storage. We propose Leopard, a hardware pass-through remote storage access framework with queue concurrency, which provides lossless remote storage access for edge intelligence. Leopard proposes a custom NVMe controller using SmartNIC’s FPGA core to emulate it as an NVMe device, which eliminates complex remote storage stacks for edge workstations. Operations for remote storage access are implemented as hardware circuits inside the controller to eliminate CPU cycles. Parallelized and pipelined workflows are proposed for hardware circuits to accelerate remote storage access operations. Our evaluation presents that Leopard exhibits $1.09 \times \sim 6.04 \times$ lower remote storage access latency than SOTA solutions for realistic workloads in edge intelligent workstations.
Bo Peng 0043, Jianguo Yao 0002, Haibing Guan
DAC2
2025 ReHSS: Optimizing Latency for Cloud Hybrid Storage Systems Using in-Network Placement
abstract
Modern cloud hybrid storage systems have been concentrating on strategic data placement to provide low disk I/O latency for various workloads. However, previous studies primarily focus on exploring adaptive data placement algorithms with high placement accuracy, overlooking the computing latency introduced by these algorithms. It significantly increases the end-to-end latency that determines the quality of service (QoS) for hybrid storage systems. We propose ReHSS, a novel in-network data placement framework that optimizes the end-to-end latency for hybrid storage systems using modern SmartNIC. We first investigate the overhead of adaptive data placement for hybrid storage systems. Then, we propose a comprehensive hardware/software co-optimization solution based on in-network processing that includes algorithm acceleration, data transmission and processing, and computing and communication overlapping. Experimental results present that compared to the SOTA solution, ReHSS optimizes the end-to-end latency for hybrid storage systems by$1.54 \times \sim 15.18 \times$.
Bo Peng 0043, Jianguo Yao 0002, Haibing Guan
IWQoS2
2025 <tt>STRCMP</tt>: Integrating Graph Structural Priors with Language Models for Combinatorial Optimization
Xijun Li, Jiexiang Yang, Bo Peng 0043, Jianguo Yao 0002, Haibing Guan
NeurIPS4
2025 SHC-DP: Software-hardware collaborative in-network data placement for hybrid storage systems
Bo Peng 0043, Jianguo Yao 0002, Haibing Guan
J. Syst. Archit.2
2024 DICFaaS: Parallel and Scalable Data Integrity Computation for Serverless Systems
abstract
Serverless computing is extensively employed in distributed storage computing scenarios owing to its dynamic responsiveness and on-demand allocation capabilities. Within distributed storage computing, data integrity computation holds significant importance. Nevertheless, current serverless systems fail to consider the CPU-intensive nature of data integrity computation, leading to prolonged CPU occupancy, increased latency, and substantial performance deterioration, particularly in parallel scenarios. Therefore, there is an imperative need for a solution that can efficiently handle this kind of computation in serverless systems. In this paper, we introduce DICFaaS, a high-performance serverless system specialized for parallel and scalable data integrity computation. DICFaaS can recognize the offloading potential of servers and automatically offload CPU-intensive data integrity computation. The offloading function of DICFaaS can be easily integrated into other existing serverless systems. We evaluate the effectiveness of DICFaaS for data integrity computation compared to state-of-the-art serverless systems. For example, DICFaaS delivers up to $17 \times$ latency improvement over AWS Lambda and up to $42 \times$ latency improvement over OpenFaaS when the data size is 200M. It also saves up to $\mathbf{6 0 \%}$ CPU overhead at runtime compared to OpenFaaS.
Tianchen Xiong, Bo Peng 0043
ICPADS2
2023 LPNS: Scalable and Latency-Predictable Local Storage Virtualization for Unpredictable NVMe SSDs in Clouds
Bo Peng 0043, Jianguo Yao 0002, Haibing Guan
USENIX ATC1
2023 FlexHM: A Practical System for Heterogeneous Memory with Flexible and Efficient Performance Optimizations
abstract
With the rapid development of cloud computing, numerous cloud services, containers, and virtual machines have been bringing tremendous demands on high-performance memory resources to modern data centers. Heterogeneous memory, especially the newly released Optane memory, offer appropriate alternatives against DRAM in clouds with the advantages of larger capacity, lower purchase cost, and promising performance. However, cloud services suffer serious implementation inconvenience and performance degradation when using hybrid DRAM and Optane memory. This article proposes FlexHM, a practical system to manage transparent heterogeneous memory resources and flexibly optimize memory access performance for all VMs, containers, and native applications. We present an open-source prototype of FlexHM in Linux with several main contributions. First, FlexHM raises a novel two-level NUMA design to manage DRAM and Optane memory as transparent main memory resources. Second, FlexHM provides flexible and efficient memory management, helping optimize memory access performance or save purchase costs of memory resources for differential cloud services with customized management strategies. Finally, the evaluations show that cloud workloads using 50% Optane slow memory on FlexHM can achieve up to 93% of the performance when using all-DRAM, and FlexHM provides up to 5.8× improvement over the previous heterogeneous memory system solution when workloads use the same ratio of DRAM and Optane memory.
Bo Peng 0043, Yaozu Dong, Jianguo Yao 0002, Fengguang Wu, Haibing Guan
ACM Trans. Archit. Code Optim.1
2022 MDev-NVMe: Mediated Pass-Through NVMe Virtualization Solution With Adaptive Polling
abstract
The fast access to data and high parallel processing in high-performance computing instigates an urgent demand on the improvement of the NVMe storage within modern data centers. However, the former NVMe virtualization’s unsatisfactory performance demonstrates that NVMe devices are often underutilized within cloud platforms. An NVMe virtualization mechanism with high performance and device sharing has captured researchers and developers’ attention. This article introduces MDev-NVMe, a new virtualization solution for NVMe storage device with (1) full NVMe storage virtualization for VMs running native NVMe driver, (2) a mediated pass-through mechanism for NVMe management, and (3) adaptive configuration of active polling optimization to simultaneously achieve high throughput, low latency performance, and substantial device scalability. We practically implement the MDev-NVMe as a Linux kernel module. This article subsequently evaluates MDev-NVMe with Intel OPTANE and P3600 SSD by comparing several mainstream NVMe virtualization mechanisms using application-level I/O benchmarks. MDev-NVMe with active polling can demonstrate a 142 percent improvement over native (interrupt-driven) throughput and over 2.5 × theVirtiothroughput with only 70 percent native average latency and 31 percentVirtioaverage latency. Finally, the advantages of MDev-NVMe and the importance of adaptive polling are discussed, offering evidence that MDev-NVMe is a superior virtualization choice for cloud storage.
Bo Peng 0043, Jianguo Yao 0002, Yaozu Dong, Haibing Guan
IEEE Trans. Computers1
2021 A Throughput-Oriented NVMe Storage Virtualization With Workload-Aware Management
abstract
Storage virtualization is an important component of large-scale online services in multi-tenant clouds. It typically shares the physical storage among guest machines and performs transactional operations for high-performance data processing. However, even with the recent mediated pass-through virtualization optimization, the operations of multi-tenant storage I/O meet the bottleneck, and thus degrade the throughput performance of the cloud storage services. We observe that the root cause of the problem is the unawareness of varying and imbalanced workload inefficiency of resource management in the multi-tenant cloud storage setting. In this paper, we present FinNVMe, a new throughput-oriented NVMe storage virtualization management mechanism, that (1) passes-through I/O performance-critical resources and emulates privileged resources to provide high throughput in a workload-aware manner among multi-tenant VMs, (2) enables fine-grained scheduling for I/O resources to achieve promising flexibility and scalability with respective to virtualization, and (3) adopts the queue binding and the queue shuffling to reduce the virtualization and management overhead, and involves active polling for further I/O acceleration. This article subsequently evaluates FinNVMe with micro benchmarks on two typical scenarios (both balanced and imbalanced workload) and the real-world storage workloads to show its high throughput performance, along with the flexibility and scalability of virtualization and resource management. For example, FinNVMe achieves up to 20 percent throughput improvement with more stable latency in the varying and imbalanced workload.
Bo Peng 0043, Ming Yang 0021, Jianguo Yao 0002, Haibing Guan
IEEE Trans. Computers1
2019 Proactive coordination for low-congestion multi-path datacenter networks
Bo Peng 0043, Jianguo Yao 0002, Haibing Guan
J. Syst. Archit.1
2018 HybridPass: Hybrid Scheduling for Mixed Flows in Datacenter Networks
abstract
Modern cloud applications generate millions of mixed flows transmitted between distributed nodes, and the typical latency-sensitive and throughput-intensive flows coexist in datacenter networks. Scheduling those mixed flows presents new challenges when meeting both low latency and high throughput requirements. This paper introduces HybridPass, a novel hybrid network architecture for datacenter networks, which is the first attempt to support both time-triggered and event-triggered scheduling in respect to latency-sensitive and throughput-intensive flows. To this end, we develop an arbiter, which uses a loosely synchronized time-triggered manner to allocate the network bandwidth for latency-sensitive and throughput-intensive flows from a global perspective. Specifically, the time-triggered scheduling aims to minimize the latency through establishing flow-level and task-level models. Then the event-triggered scheduling is developed to utilize the leftover bandwidth for throughput-intensive flows without any impact on latency-sensitive flows. Our experiments show that HybridPass can achieve up to 40.71% latency reduction for latency-sensitive flows compared with the baseline DCTCP while maintaining the high throughput with inappreciable 0.77% throughput sacrifice for throughput-intensive flows.
Bo Peng 0043, Jianguo Yao 0002, Zhengwei Qi, Haibing Guan
IPDPS1
2018 MDev-NVMe: A NVMe Storage Virtualization Solution with Mediated Pass-Through
Bo Peng 0043, Haozhong Zhang, Jianguo Yao 0002, Yaozu Dong, Haibing Guan
USENIX ATC1