VLDB 2026 Research / reviewers in the wild / expert
Yiren Pang
dblp:339/7289
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2024
0009-0001-3296-0378ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Cloud and datacenter computing · 60% High-performance computing · 20% Distributed systems · 20% | |
| Artificial intelligence
1 paper |
Generative modeling · 100% | |
| Computer networks
1 paper |
Datacenter networks · 100% | |
| Computer graphics and multimedia
1 paper |
Computer animation and physical simulation · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Operating systems · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Datacenter networks
datacenter transport |
0.8 | 1 | 2024 | PPT: A Pragmatic Transport for Datacenters · SIGCOMM 2024 |
Machine learning › Generative modeling
diffusion model |
0.7 | 1 | 2023 | HumanMAC: Masked Motion Completion for Human Motion Prediction · ICCV 2023 |
Machine learning › Generative modeling › diffusion model
motion diffusion |
0.7 | 1 | 2023 | HumanMAC: Masked Motion Completion for Human Motion Prediction · ICCV 2023 |
Computer animation and physical simulation
human motion prediction |
0.7 | 1 | 2023 | HumanMAC: Masked Motion Completion for Human Motion Prediction · ICCV 2023 |
Cloud and datacenter computing › cloud networking
container networking |
0.7 | 1 | 2023 | Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023 |
Cloud and datacenter computing › cloud networking › container networking
container overlay network |
0.7 | 1 | 2023 | Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023 |
High-performance computing › data transfer
data delivery |
0.7 | 1 | 2023 | Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023 |
Distributed systems › peer-to-peer systems
overlay networks |
0.7 | 1 | 2023 | Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023 |
Cloud and datacenter computing
virtualization |
0.7 | 1 | 2023 | Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023 |
Datacenter networks
flow scheduling |
0.2 | 1 | 2024 | PPT: A Pragmatic Transport for Datacenters · SIGCOMM 2024 |
Operating systems
network stack |
0.2 | 1 | 2023 | Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023 |
Operating systems › operating system interface
system call |
0.2 | 1 | 2023 | Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay Network · IEEE Trans. Parallel Distributed Syst. 2023 |
Methods — techniques the papers use, named apart from their topics
syscall threshold · 1.3multithreading · 1.3masked completion · 1.3denoising diffusion · 1.3intermittent loop initialization · 0.8exponential window decrease · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | WQEFC: A Scalable and Low-Latency RDMA Messages Scheduler for Mixed MessagesabstractRDMA has been widely deployed to improve the performance of applications with frequently fine-grained remote access. However, restricted on-chip resources result in cache misses under high concurrency that significantly degrade network performance. QPC-aware solutions only focus on the number of concurrent QPs, ignoring the impact of WQE within QPs. SMART limits the number of WQEs in each QP with a credit-based scheme. Nevertheless, we find that equal treatment increases the tail latency of messages ranging from 32 bytes to 1024 bytes by 2× when mixing messages of different sizes. In this paper, we introduce WQEFC, a scalable RDMA message scheduler that provides lower latency and higher throughput for applications with heavily concurrent messages. Our key insight is that there is a significant difference in the sensitivity to cache miss between messages of different sizes. For messages smaller than 32 bytes, which are sensitive to cache misses, we combine the credit limiter and sub-message poller, limiting the number of concurrent wqes to avoid cache miss while ensuring optimal message completion latency. For other messages, which are insensitive to cache miss, we assign them a higher priority and use the sub-message poller to ensure message concurrency while reducing cache miss. We implement WQEFC as a middleware between the driver layer and the application layer for flexible deployment. WQEFC outperforms the state-of-the-art solution Smart by increasing system throughput by 61.3%, and reducing the tail latency of messages smaller than 32 bytes and larger than 32 bytes by 41.6% and 72.6%, respectively. Yaozhen Li, Lide Suo, Xiancheng Meng, Yiren Pang, Wenxin Li 0001, Keqiu Li, Yitao Hu |
HPCC | 5 |
| 2024 | PPT: A Pragmatic Transport for DatacentersabstractThis paper introduces PPT, a pragmatic transport that achieves comparable performance to proactive transports while maintaining good deployability as reactive transports. Our key idea is to run a low-priority control loop to leverage the available bandwidth left by the reactive transports. The main challenge is to send just enough packets to improve performance without harming the primary control loop. We combine two unconventional techniques: an intermittent loop initialization and an exponential window decrease, enabling us to dynamically identify and fill the spare bandwidth. We further complement PPT's design with a buffer-aware flow scheduling scheme to optimize the average FCT of small flows without prior knowledge of flow size information. We have implemented a PPT prototype in the Linux kernel with ~400 lines of code and demonstrated that compared to Homa, it delivers up to 46.3% lower overall average FCT and even 25%/55.5% lower average/tail FCT of small flows in an Memcached workload. Lide Suo, Yiren Pang, Wenxin Li 0001, Renjie Pei, Keqiu Li, Xiulong Liu 0001, Xin He 0043, Yitao Hu, Guyue Liu |
SIGCOMM | 2 |
| 2023 | HumanMAC: Masked Motion Completion for Human Motion PredictionabstractHuman motion prediction is a classical problem in computer vision and computer graphics, which has a wide range of practical applications. Previous effects achieve great empirical performance based on an encoding-decoding style. The methods of this style work by first encoding previous motions to latent representations and then decoding the latent representations into predicted motions. However, in practice, they are still unsatisfactory due to several issues, including complicated loss constraints, cumbersome training processes, and scarce switch of different categories of motions in prediction. In this paper, to address the above issues, we jump out of the foregoing style and propose a novel framework from a new perspective. Specifically, our framework works in a masked completion fashion. In the training stage, we learn a motion diffusion model that generates motions from random noise. In the inference stage, with a denoising procedure, we make motion prediction conditioning on observed motions to output more continuous and controllable predictions. The proposed framework enjoys promising algorithmic properties, which only needs one loss in optimization and is trained in an end-to-end manner. Additionally, it accomplishes the switch of different categories of motions effectively, which is significant in realistic tasks, e.g., the animation task. Comprehensive experiments on benchmarks confirm the superiority of the proposed framework. The project page is available at https://lhchen.top/Human-MAC. Yewen Li, Yiren Pang, Xiaobo Xia, Tongliang Liu |
ICCV | 4 |
| 2023 | MiddleCache: Accelerating TCP based In-memory Key-value Stores using eBPFabstractIn-memory key-value stores are widely used in modern web services to support large-scale user requests by caching popular data. Their performance is critical, and BMC, the state-of-the-art work, builds an in-kernel cache and processes requests before the stack using eBPF to reduce the overhead of the kernel network stack. However, BMC fails to support stateful protocol TCP because pre-stack processing creates TCP state bias between the client and server.TCP is widely used by in-memory key-value stores, is even the only choice for some applications (e.g., Redis), and also suffers from performance issues. In this work, we present MiddleCache, a TCP-enabled in-memory key-value store acceleration design. Our key observation is that the TCP state bias of the client and server can be inferred and eliminated with packet length. The design of MiddleCache has two key parts: (i) A compact TCP state maintenance mechanism that accumulates packet lengths and applies corrections to the packet header, which realize TCP support within the constrains of eBPF. (ii) Lock-free accumulation counters that support high-performance concurrent access by utilizing Receive Side Scaling (RSS). Our experiments show that, compared with Memcached, MiddleCache reduces 56% processing latency on cache hit and achieves a 3.8× throughput improvement on Facebook-like small-size requests workload. Yiren Pang, Sheng Chen 0015, Wenxin Li 0001, Yulong Li 0001, Xin He 0043, Song Zhang 0008, Zewei Guan, Lide Suo |
ICPADS | 1 |
| 2023 | Accelerating Data Delivery of Latency-Sensitive Applications in Container Overlay NetworkabstractContainer overlay network, though being widely adopted to enable communication between containers on different hosts, is a key downside for latency-sensitive applications. The state-of-the-art solution seeks to shorten the data path in packet processing by replacing overlay connection file descriptors with host namespace ones. While promising, it must block each overlay connection until the relevant host connection is set up, thus heavily influencing the request latency. In this paper, we present ShuntFlow, a systematic data delivery framework that seamlessly integrates the host and overlay networks to reduce the application's request-response latency. ShuntFlow first lets all connections flow in the overlay network directly. Then, it adopts a simple-yet-effective syscall-threshold-based mechanism to pick appropriate connections and switches their data delivery to the host network in a blocking-free way using a multi-threading technique. As such, unnecessary connection switches are prevented; yet, the pre-setup phase dilemma is eliminated. We have implemented a ShuntFlow prototype based on Linux and Docker and evaluated it extensively on a 40 Gbps testbed. The results show that ShuntFlow achieves 13%/72% and 19%/69% reductions, in average/tail request-response latency of a web server and an in-memory key-value store, respectively, while incurring less CPU overhead, compared to Slim. Wenxin Li 0001, Yiren Pang, Renjie Pei, Yitao Hu, Lide Suo, Keqiu Li |
IEEE Trans. Parallel Distributed Syst. | 3 |