Renjie Pei

dblp:361/5579 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0007-4560-2311ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 42% Cloud and datacenter computing · 31% Memory systems · 16%
Computer networks
1 paper
Datacenter networks · 100%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems
concurrency control
1.012026
ECC: Efficient Concurrency Control for Disaggregated Memory Systems · IEEE Trans. Computers 2026
Distributed systems
distributed coordination
1.012026
ECC: Efficient Concurrency Control for Disaggregated Memory Systems · IEEE Trans. Computers 2026
Memory systems
memory disaggregation
1.012026
ECC: Efficient Concurrency Control for Disaggregated Memory Systems · IEEE Trans. Computers 2026
Datacenter networks
datacenter transport
0.812024
PPT: A Pragmatic Transport for Datacenters · SIGCOMM 2024
Cloud and datacenter computing › cloud networking
container networking
0.712023
Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023
Cloud and datacenter computing › cloud networking › container networking
container overlay network
0.712023
Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023
High-performance computing › data transfer
data delivery
0.712023
Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023
Distributed systems › peer-to-peer systems
overlay networks
0.712023
Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023
Cloud and datacenter computing
virtualization
0.712023
Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023
Datacenter networks
flow scheduling
0.212024
PPT: A Pragmatic Transport for Datacenters · SIGCOMM 2024
Operating systems
network stack
0.212023
Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023
Operating systems › operating system interface
system call
0.212023
Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023

Methods — techniques the papers use, named apart from their topics

syscall threshold · 1.3multithreading · 1.3retry interval prediction · 1.0optimistic polling · 1.0intermittent loop initialization · 0.8exponential window decrease · 0.8
YearPublicationVenuePosition
2026 ECC: Efficient Concurrency Control for Disaggregated Memory Systems
abstract
Memory disaggregation architecture presents unique challenges in ensuring data consistency due to the limited computational power available at memory servers. One line of pessimistic solutions utilizes lock tables to deal with this challenge. Since the lock release signal cannot be immediately synchronized to the compute server, such solutions can only achieve suboptimal memory utilization. To address the aforementioned limitation, another line of solutions operates optimistically, polling the memory server until the operation completes successfully. However, these solutions trigger a massive number of unnecessary retries, resulting in performance collapse in high contention scenarios.high memory utilization while avoiding the need for retries. Our key idea is that compute servers proactively predict the appropriate retry interval and schedule accordingly. By analyzing the properties of requests on memory disaggregation, we find that the retry interval mainly consists of network fluctuations and host congestion delay. ECC design builds upon insights gained from this analysis. We have integrated ECC into Clover—a state-of-the-art memory disaggregation system and evaluated it through simulations and testbed experiments. Our testbed results show that ECC improves throughput by 50.8% over Clover.
Yang Li 0245, Yaozhen Li, Wenxin Li 0001, Yulong Li 0001, Song Zhang 0008, Renjie Pei, Xiancheng Meng, Keqiu Li
IEEE Trans. Computers6
2024 Mild: A Zero-Wait Multi-Round Proactive Transport
abstract
With the rapid growth of datacenter network link speed, multi-round matching based proactive solutions (e.g., dcPIM) has become increasingly attractive. Such solutions enable receivers to obtain as much global information as possible through multi-round matching, thereby facilitating them to make near-optimal decisions on bandwidth allocation. However, the matching phase before transmitting data introduces significant latency overhead. In this paper, we present Mild, a zero-wait solution that runs a second sender-driven control loop in parallel, leveraging in-network telemetry (INT) to detect and fill the spare bandwidth during the matching phase. Furthermore, we introduce a selective dropping mechanism to ensure that the packets from the second loop do not impact the data transmission of the primary loop. Additionally, we use the well-protected primary loop to perform loss recovery for the dropped packets efficiently. We integrate Mild into a representative proposal dcPIM and evaluate its performance through 100Gbps large-scale simulations. Compared to the state-of-the-art solution, Mild reduces the tail flow completion time (FCT) of short flows by up to 55% while achieving up to 57%/45% lower average FCT of medium/large flows.
Renjie Pei, Wenxin Li 0001, Yulong Li 0001, Song Zhang 0008, Yaozhen Li, Wenyu Qu
ISCC1
2024 PPT: A Pragmatic Transport for Datacenters
abstract
This paper introduces PPT, a pragmatic transport that achieves comparable performance to proactive transports while maintaining good deployability as reactive transports. Our key idea is to run a low-priority control loop to leverage the available bandwidth left by the reactive transports. The main challenge is to send just enough packets to improve performance without harming the primary control loop. We combine two unconventional techniques: an intermittent loop initialization and an exponential window decrease, enabling us to dynamically identify and fill the spare bandwidth. We further complement PPT's design with a buffer-aware flow scheduling scheme to optimize the average FCT of small flows without prior knowledge of flow size information. We have implemented a PPT prototype in the Linux kernel with ~400 lines of code and demonstrated that compared to Homa, it delivers up to 46.3% lower overall average FCT and even 25%/55.5% lower average/tail FCT of small flows in an Memcached workload.
Lide Suo, Yiren Pang, Wenxin Li 0001, Renjie Pei, Keqiu Li, Xiulong Liu 0001, Xin He 0043, Yitao Hu, Guyue Liu
SIGCOMM4
2023 Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network
abstract
Container overlay network, though being widely adopted to enable communication between containers on different hosts, is a key downside for latency-sensitive applications. The state-of-the-art solution seeks to shorten the data path in packet processing by replacing overlay connection file descriptors with host namespace ones. While promising, it must block each overlay connection until the relevant host connection is set up, thus heavily influencing the request latency. In this paper, we present ShuntFlow, a systematic data delivery framework that seamlessly integrates the host and overlay networks to reduce the application's request-response latency. ShuntFlow first lets all connections flow in the overlay network directly. Then, it adopts a simple-yet-effective syscall-threshold-based mechanism to pick appropriate connections and switches their data delivery to the host network in a blocking-free way using a multi-threading technique. As such, unnecessary connection switches are prevented; yet, the pre-setup phase dilemma is eliminated. We have implemented a ShuntFlow prototype based on Linux and Docker and evaluated it extensively on a 40 Gbps testbed. The results show that ShuntFlow achieves 13%/72% and 19%/69% reductions, in average/tail request-response latency of a web server and an in-memory key-value store, respectively, while incurring less CPU overhead, compared to Slim.
Wenxin Li 0001, Yiren Pang, Renjie Pei, Yitao Hu, Lide Suo, Keqiu Li
IEEE Trans. Parallel Distributed Syst.4