VLDB 2026 Research / reviewers in the wild / expert
Ruyi Yao
dblp:273/1828
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-5875-1953ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 14 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Themis: Scheduling-Aware Buffer Management for HBM-Based Hybrid Buffers
Zhiyu Zhang 0012, Minkun Xue, Ruyi Yao, Shili Chen, Yibo Fan, Yang Xu 0010 |
NSDI | 4 |
| 2026 | Scale-up PIFO: Interleaving Multiple Priority Queues for High Speed Programmable SchedulingabstractPush-In First-Out (PIFO) offers a unified abstraction for rapidly deploying diverse scheduling algorithms on the same hardware. As SerDes-lane aggregation pushes port rates to 1.6 Tbps, the perpacket processing budget is at a sub-nanosecond scale, making single-queue PIFO designs fail to keep up. Mirroring lane aggregation, we advocate interleaving multiple PIFO queues. However, simple round-robin parallelization introduces substantial scheduling error, and in the worst case, it can grow to the order of the buffer size. Shili Chen, Ruyi Yao, Zhiyu Zhang 0012, Hao Wang 0231, Deli Huang, Yibo Fan, Yang Xu 0010 |
SIGCOMM | 5 |
| 2026 | InfiniFlow: Decoupling Virtual Channel Scalability from Buffer Requirements in Lossless Datacenter Networks
Zerui Tian, Sen Liu 0002, Minkun Xue, Hao Shangguan, Ruyi Yao, Deli Huang, Songchen Xue, Yang Xu 0010 |
SIGCOMM | 5 |
| 2025 | Empowering Flowlet Load Balancing in RDMA with Host-Based Flowlet Fine-TuningabstractFlowlet-level load balancing has not demonstrated the expected robust capability in RDMA networks due to insufficient flowlets and the adverse effects of PFC. To delve deeper, we conduct measurements at end hosts and perform a detailed analysis of time gaps between packets. Our investigation reveals that in RDMA networks, the number of time gaps exceeding the flowlet timeout is considerably lower than the number in TCP networks. We also identify a stepwise time gap pattern that predicts the occurrence of PFC. Based on these observations, we propose$\text{HF}^{2} \mathrm{T}$, a host-based time gap adjustment method to improve the effectiveness of flowlet-level load balancing in RDMA networks. The core idea involves delaying a minimal number of specific packets at the host, actively extending the time gaps between them, thereby fostering the generation of sufficient flowlets at the switch and enhancing the utilization of equal-cost links. Incorporating an identification algorithm for the time gap pattern that predicts PFC,$\text{HF}^{2} \mathrm{T}$also leverages the time gap extension to reroute traffic away from potential PFC paths in advance, thus mitigating PFC occurrences. The minor cost of delaying a few packets is vastly offset by the benefits of generating flowlets and reducing PFC. We use DPDK to implement a prototype of$\text{HF}^{2} \mathrm{T}$, and through testbed experiments, we demonstrate that$\text{HF}^{2} \mathrm{T}$, serving as a building block for flowlet load balancing, can enhance the throughput of CONGA by 16.82%. The simulation results also show that$\text{HF}^{2} \mathrm{T}$can reduce the average FCT by 16.38% and the 99-percentile FCT by 21.13% compared to the state-of-the-art RDMA load balancing ConWeave. Chuhao Chen 0001, Deli Huang, Zerui Tian, Ruyi Yao, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 4 |
| 2025 | CClinguist: An Expert-Free Framework for Future-Compatible Congestion Control Algorithm IdentificationabstractCongestion control algorithms (CCAs) play a critical role in determining transmission quality. With their rapid evolution during the past few decades, understanding the CCA landscape on the Internet has become increasingly essential for network advancement. Traditional CCA census tools, however, rely heavily on manual configuration and construction, necessitating significant human effort to keep pace with the introduction of new CCAs. Ruyi Yao, Jialin Wei, Ruoshi Sun, Sen Liu 0002, Yang Xu 0010 |
SIGCOMM | 3 |
| 2024 | TCAMVisor: High-throughput TCAM Virtualization for Multi-tenant Software Defined NetworkingabstractSoftware Defined Networking (SDN) provides users with a unified abstraction of physical networks. To meet the demands of modern data centers, many works have focused on designing network virtualization hypervisors that support multi-tenant SDN. Ternary Content Addressable Memory (TCAM) is widely used in SDN switches for rule storage. While it has extremely high lookup throughput, it also features drawbacks such as small capacity and slow update speed. Faced with multi-tenant scenarios, its limitations are even more pronounced. Existing hypervisors lack consideration for TCAM isolation, leading to slower TCAM updates and the mutual impact of requests from different tenants. Consequently, they fail to provide guaranteed performance to tenants. To solve these problems, we propose TCAMVisor, which further isolates TCAM resources based on traditional SDN hypervisors. Specifically, TCAMVisor provides better allocation mechanisms for TCAM entry and control bandwidth, ensuring inter-tenant isolation while improving resource utilization. Additionally, TCAMVisor improves the update speed of TCAM by delicately placing tenant rules. To the best of our knowledge, TCAMVisor is the first work to effectively achieve tenant isolation in TCAM, with an average throughput improvement of 5.5 times compared to FlowVisor. Ruoshi Sun, Ruyi Yao, Hao Wang 0231, Yiren Zhou, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 3 |
| 2024 | vPIFO: Virtualized Packet Scheduler for Programmable Hierarchical Scheduling in High-Speed NetworksabstractProgrammable packet scheduling enables the integration of scheduling algorithms into switches without the need for hardware redesign. The Push-In First-Out (PIFO) queue facilitates a programmable packet scheduler, supporting a single scheduling algorithm flexibly. However, hierarchical scheduling required in Multi-Tenant Data Centers (MTDCs) remains non-programmable. Dynamic and diverse hierarchical scheduling algorithms necessitate alterations in both the number of PIFO queues and their connection topology, posing a significant challenge to support them on fixed hardware. Zhiyu Zhang 0012, Shili Chen, Ruyi Yao, Ruoshi Sun, Hao Wang 0231, Gaojian Fang, Yibo Fan, Wanxin Shi, Sen Liu 0002, Yang Xu 0010 |
SIGCOMM | 3 |
| 2023 | CoLUE: Collaborative TCAM Update in SDN Switches
Ruyi Yao, Chuhao Chen 0001, Wenjun Li 0004, Ying Wan 0001, Sen Liu 0002, Bin Liu 0001, Yang Xu 0010 |
INFOCOM | 1 |
| 2023 | Rusen: Rule Semantics Enabler toward Fast TCAM Update for Commodity SDN SwitchesabstractTernary Content Addressable Memory (TCAM) is widely used in Software-Defined Networking (SDN) switches due to its impressive throughput. But its unique circuit design results in long and inconsistent update delays. To overcome this challenge, many TCAM update algorithms based on rule semantics have been proposed. These algorithms eliminate unnecessary order restrictions, thus reducing update delays in theory. However, most commodity switches are semantic-unaware, which maintain rules in strict priority order. These algorithms are therefore not available for practical use. To address this issue, this paper proposes Rusen, a framework that enables the use of many semantic-based algorithms on Semantic-unaware commodity switches. Working as a transparent middle layer, the core idea of Rusen is to express the update scheme derived by semantic-based algorithms as messages that the Semantic-unaware switches can execute. In addition, Rusen optimizes the update scheme based on the specific characteristics of each switch, leading to improved performance of these algorithms. We evaluate the performance of Rusen by enabling several state-of-the-art semantic-based algorithms on commodity SDN switches. Results show that the average update delay can be significantly reduced by 23%∼94% on OpenFlow switches and 39%∼84% on a P4 switch. Ruoshi Sun, Ruyi Yao, Chuhao Chen 0001, Sen Liu 0002, Yang Xu 0010 |
IWQoS | 3 |
| 2023 | BMW Tree: Large-scale, High-throughput and Modular PIFO Implementation using Balanced Multi-Way Sorting TreeabstractPush-In-First-Out (PIFO) queue has been extensively studied as a programmable scheduler. To achieve accurate, large-scale, and high-throughput PIFO implementation, we propose the Balanced Multi-way (BMW) Sorting Tree for real-time packet sorting. The tree is highly modularized, insertion-balanced and pipeline-friendly with autonomous nodes. Ruyi Yao, Zhiyu Zhang 0012, Gaojian Fang, Peixuan Gao, Sen Liu 0002, Yibo Fan, Yang Xu 0010, H. Jonathan Chao |
SIGCOMM | 1 |
| 2022 | BubbleTCAM: Bubble Reservation in SDN Switches for Fast TCAM UpdateabstractThe unique hardware structure of Ternary Content-Addressable Memory (TCAM) enables its unparalleled lookup throughput but also causes slow update due to the Priority Order Constraint (POC). With the increase of application demands, TCAM update has become a bottleneck in the network. This paper proposes a new TCAM management mechanism named BubbleTCAM to enable fast TCAM update, in which available empty entries are defined as bubbles. The core idea of Bub-bleTCAM is to uniformly distribute bubbles and dependency chains in TCAM, which is beneficial to updates. BubbleTCAM consists of two components: bubble management and rule insertion. Bubble management enables TCAM to have uniformly distributed bubbles at all times through three key procedures: bubble lock reservation, bubble lock release and bubble generation. Rule insertion ensures that dependency chains of rules are uniformly stretched and distributed in TCAM. In addition, BubbleTCAM avoids the reorder problem by pre-sorting. Our evaluation based on the rulesets generated by ClassBench shows that BubbleTCAM effectively reduces the average cost and worst cost (in units of rule movements) during rule updates by at least 48% and 50%, respectively. Especially for the worst cost, the performance can be improved by up to 196x. Chuhao Chen 0001, Ruyi Yao, Ying Wan 0001, Wenjun Li 0004, Sen Liu 0002, Bin Liu 0001, Yang Xu 0010 |
IWQoS | 4 |
| 2021 | KickTree: A Recursive Algorithmic Scheme for Packet Classification with Bounded Worst-Case PerformanceabstractAs a promising alternative to TCAM-based solutions for packet classification, FPGA has received increasing attention. Although extensive research has been conducted in this area, existing FPGA-based packet classifiers cannot satisfy the burgeoning needs from OpenFlow, which demands large-scale rule sets and frequent rule updates. As a recently proposed hardware-specific approach, TabTree avoids rule replication and supports dynamic rule update. However, it still faces problems of unbalanced rule subset partition, unevenly distributed subtrees and excessive TSS leaf nodes when implemented on FPGA. In this paper, we propose a hardware-friendly packet classification approach called KickTree, which is elaborated by considering hardware properties. To take advantage of intrinsic parallelism of FPGA, KickTree adopts multiple balanced decision trees which can run simultaneously. The bit selection is more flexible which breaks the restriction of rule subset. Moreover, each subset size is strictly limited, leading to bounded and evenly-distributed Yao Xin, Yuxi Liu 0017, Wenjun Li 0004, Ruyi Yao, Yang Xu 0010, Yi Wang 0004 |
ANCS | 4 |
| 2021 | MagicTCAM: A Multiple-TCAM Scheme for Fast TCAM UpdateabstractTernary Content-Addressable Memory (TCAM) is a popular solution for high-speed flow table lookup in Software-Defined Networking (SDN). Rule insertion in TCAM is a time-consuming operation. To ensure semantic correctness, rules overlapped must be stored in TCAM with decreasing priority order and many rule movements may be needed to make space for a single inserted rule. When a rule insertion is in progress, the regular flow table lookup will be suspended, which could lead to a degraded user experience for SDN applications. In this paper, we propose a multiple-TCAM framework named MagicTCAM to reduce the rule movements during a rule insertion. The core of MagicTCAM lies in three operations: layering, partitioning and rotating. By layering, rules with the least overlapping will be grouped (i.e., layered) into a sub-ruleset. The number of rule movements is therefore greatly reduced as most of rules in a sub-ruleset are non-overlapped. To achieve balanced load in TCAMs, rules in each sub-ruleset are further partitioned and dispatched into different TCAMs in a rotating manner. In addition, an inter-TCAM movement algorithm is proposed to allow rules to be moved between TCAMs for reduced rule movement. Experiment results show that with two half-sized TCAMs, MagicTCAM reduces the rule movements by 39% on average compared with the state-of-the-art work while the computation time is shortened by half as well. Ruyi Yao, Xuandong Liu, Ying Wan 0001, Bin Liu 0001, Wenjun Li 0004, Yang Xu 0010 |
ICNP | 1 |
| 2021 | PIPO: Efficient Programmable Scheduling for Time Sensitive NetworkingabstractTime Sensitive Networking (TSN) is an emerging Ethernet technology for real-time systems. To address different Quality-of-Service (QoS) requirements of applications, IEEE 802.1 TSN Task Group has standardized several packet scheduling and shaping algorithms. The software implementation of these algorithms is hard to meet the performance requirements, while the hardware implementation in Application-Specific Integrated Circuit (ASIC) is inflexible. A hardware-programmable scheduler is necessary to deal with this dilemma. Among the existing primitives, the most expressive one is Push-In-Extract-Out (PIEO), but its complexity makes the implementation very expensive. A relatively lower-cost implementation of PIEO cannot guarantee the scheduling correctness for the most critical Time-Triggered (TT) traffic in TSN. As a remedy, in this paper we propose a new Push-In-Pick-Out (PIPO) primitive under a TSN programmable scheduling framework. Composed of simple priority queues, PIPO can express all existing TSN scheduling and shaping algorithms, and is flexible enough to support future ones. Our PIPO implementation guarantees the TT traffic scheduling correctness. The simulation results corroborate the theoretical analysis that the low-cost PIPO can closely approximate PIEO and sustain a high bandwidth utilization. The prototype on Xilinx FPGA shows that, with 2,048 inputs, the PIPO-based scheduler achieves a throughput of 70 Mpps, which is 1.64x higher than the PIEO-based one, but using only 14.7% Look-Up Tables (LUTs) and 40.5% Block RAMs of the latter. Chuwen Zhang, Zhikang Chen, Haoyu Song 0001, Ruyi Yao, Yang Xu 0010, Yi Wang 0004, Ji Miao, Bin Liu 0001 |
ICNP | 4 |
| 2021 | Routing optimization meets Machine Intelligence: A perspective for the future network
Bin Dai 0002, Yuanyuan Cao, Zhongli Wu, Zhewei Dai, Ruyi Yao, Yang Xu 0010 |
Neurocomputing | 5 |