VLDB 2026 Research / reviewers in the wild / expert
Biao Lyu
dblp:267/2853 · also Biao Lyv
· DBLP profile ↗
39ranked-venue papers
1as first author
37since 2021 · last 2026
0000-0002-8096-5528ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 28 · 1 first-author · 27 since 2021Systems, architecture and hardware · 7 · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Scaling LLM Agent Tool Access at Cloud ScaleabstractLLM agents increasingly rely on tool calling, and the Model Context Protocol (MCP) standardizes it between agents and tool providers, reducing integration cost and driving rapid growth in tool scale. Yet a standardized interface does not make tool access work at production scale: legacy services are not MCP-callable, fast protocol evolution creates compatibility cost, large tool sets exhaust the context window, and stateful sessions complicate load balancing. We solve these with a shared control point, a centralized MCP Gateway System that makes MCP operational at cloud scale. The gateway breaks the direct-connect data plane and consolidates legacy API integration, protocol bridging, access control, and session-aware routing, while scaling out elastically at low per-call overhead. It scales agent tool access to thousands of cloud operations. Enge Song, Yueshang Zuo, Rong Wen, Jing Tie, Zhou Shao, Qiang Fu 0011, Xiaobo Xue, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Tian Pan 0001, Yang Song 0031, Xing Li 0007, Biao Lyu, Meng Li 0010, Haipeng Dai 0001, Guihai Chen, Shunmin Zhu |
APNet | 25 |
| 2026 | SOPSmith: Forging Executable SOPs for LLM-Driven GPU Cluster Network Diagnosis
Guoyao Yu, Xiaoqing Sun, Yangyang Shi, Yang Song 0031, Xing Li 0007, Biao Lyu, Zhenguang Liu, Qinming He |
IWQoS | 8 |
| 2026 | Bifrost: Alibaba's Next-Generation VPC Network with High-Performance Multipath Reliable Transport
Xing Li 0007, Bo Jiang 0003, Yilong Lv, Yuke Hong, Yinian Zhou, Junnan Cai, Jiayue Xu, Yunrui Hu, Zhao Gao, Enge Song, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Changgang Zheng, Yang Song 0031, Biao Lyu, Rong Wen, Zhigang Zong, Shunmin Zhu |
NSDI | 25 |
| 2026 | CStar Gateway: Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Tian Pan 0001, Jin Ke 0005, Baohai Hu, Changgang Zheng, Enge Song, Donglin Lai, Yisong Qiao, Bengbeng Xue, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Yang Song 0031, Xionglie Wei, Biao Lyu, Rong Wen, Zhigang Zong, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu |
NSDI | 21 |
| 2026 | ZooRoute: Enhancing Cloud-Scale Network Reliability via Candidate Path Provisioning and Overlay Proactive Rerouting
Xiaoqing Sun, Xing Li 0007, Xionglie Wei, Tian Pan 0001, Yi Wang 0004, Chenhao Jia, Zhanlong Zhang, Xiaobo Xue, Jianyuan Lu, Shize Zhang, Enge Song, Yang Song 0031, Rong Wen, Biao Lyu, Yang Xu 0010, Shunmin Zhu |
NSDI | 20 |
| 2025 | Understanding the Long Tail Latency of TCP in Large-Scale Cloud Networks
Enge Song, Bo Jiang 0003, Yang Song 0031, Yuke Hong, Yilong Lv, Yinian Zhou, Junnan Cai, Chao Wang 0128, Yi Wang 0004, Yehao Feng, Shize Zhang, Xiaoqing Sun, Jianyuan Lu, Xing Li 0007, Biao Lyu, Zhigang Zong, Shunmin Zhu |
APNet | 20 |
| 2025 | Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Yang Song 0031, Tian Pan 0001, Zhigang Zong, Bengbeng Xue, Xionglie Wei, Yisong Qiao, Donglin Lai, Baohai Hu, Jin Ke 0005, Enge Song, Jianyuan Lu, Xing Li 0007, Biao Lyu, Rong Wen, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu |
APNet | 16 |
| 2025 | FlowCheck: Decoupling Checkpointing and Training of Large-Scale ModelsabstractCheckpointing is becoming a hotspot of interest in both academia and industry as the primary fault-tolerance method for large model training. However, existing checkpoint designs are tightly coupled with the training process, leading to interruptions that reduce overall training efficiency. To reduce the impact of checkpoints on training, this paper presents FlowCheck, a novel checkpointing system that decouples checkpoint operations from the training process, enabling checkpoint saving without blocking the training. Specifically, FlowCheck updates the checkpoints by extracting complete gradient information from the network traffic of normal training. FlowCheck deploys a traffic-mirroring network to support this design. To utilize mirrored traffic for checkpointing operations, two key challenges need to be addressed. First, we need to achieve precise identification and extraction of gradient packets from training traffic. Second, the transmission on the mirror link is unreliable due to its inability to trigger retransmission upon packet loss. Through two key designs: (1) packet-counting-based traffic identification, and (2) packet redundancy recovery mechanism, FlowCheck implements an efficient checkpointing system using the existing training network and solves the above two challenges. Experiments and estimations verify that FlowCheck achieves checkpoint operations with zero impact on training, and demonstrate that FlowCheck achieves over 98% effective training time under practical fault conditions. Zimeng Huang, Hao Nie, Haonan Jia, Bo Jiang 0003, Junchen Guo, Jianyuan Lu, Rong Wen, Biao Lyu, Shunmin Zhu, Xinbing Wang |
EuroSys | 8 |
| 2025 | FastIOV: Fast Startup of Passthrough Network I/O Virtualization for Secure ContainersabstractSingle Root I/O Virtualization (SR-IOV) technology has advanced in recent years and can simultaneously satisfy the network requirements of high data plane performance, high deployment density, and fast startup for applications in traditional containers. However, it falls short with secure containers, which have become the mainstream choice in multi-tenant clouds. SR-IOV requires secure containers to use passthrough I/O for higher data plane performance, which hinders the container startup performance and prevents its usage in time-sensitive tasks like serverless computing. In this paper, we advocate that the startup performance of SR-IOV enabled secure containers can be further boosted, making SR-IOV suitable for building a Container Network Interface (CNI) for secure containers. We first dissect the end-to-end concurrent startup process and identify three key bottlenecks that lead to the slow startup, including Virtual Function I/O device set management, Direct Memory Access memory mapping, and Virtual Function (VF) driver initialization. We then propose a CNI named FastIOV that addresses these bottlenecks through lock decomposition, unnecessary mapping skipping, decoupled zeroing, and asynchronous VF driver initialization. Our evaluation shows that FastIOV reduces the overhead of enabling SR-IOV for secure containers by 96.1%, achieving 65.7% and 75.4% reductions in the average and 99th percentile end-to-end startup time. Yunzhuo Liu, Junchen Guo, Bo Jiang 0003, Yang Song 0031, Rong Wen, Biao Lyu, Shunmin Zhu, Xinbing Wang |
EuroSys | 7 |
| 2025 | Hermes: Enhancing Layer-7 Cloud Load Balancers with Userspace-Directed I/O Event NotificationabstractLayer-7 load balancers (L7 LBs) improve service performance, availability, and scalability in public clouds. They rely on I/O event notification mechanisms such as epoll to dispatch connections from the kernel to userspace workers. However, early epoll versions suffered from the thundering herd problem. Epoll exclusive (available since Linux 4.5) mitigates this but introduces LIFO wakeups, causing connection concentration on a few workers. Reuseport (Linux 3.9) hashes connections across workers but suffers from hash collisions and lacks awareness of worker load. Since each worker serves multi-tenant traffic, inter-worker load balancing is critical to avoid worker overload and preserve tenant performance isolation. Tian Pan 0001, Enge Song, Yueshang Zuo, Shaokai Zhang, Yang Song 0031, Jiangu Zhao, Wengang Hou, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Jiao Zhang 0002, Tao Huang 0005, Biao Lyu, Xing Li 0007, Rong Wen, Zhigang Zong, Shunmin Zhu |
SIGCOMM | 14 |
| 2025 | Nezha: SmartNIC-based Virtual Switch Load SharingabstractCloud providers use SmartNIC-accelerated virtual switches (vSwitches) to offer rich network functions (NFs) for tenant VMs. Constrained by limited SmartNIC resources, it is a challenge to provide sufficient network performance for high-demand VMs. Meanwhile, we observed a significant number of idle vSwitches in the data center, which led us to consider leveraging them to build a remote resource pool for high-demand virtual NICs (vNICs). In this work, we propose Nezha, a distributed vSwitch load sharing system. Nezha reuses the existing idle SmartNICs to handle the excess load from the local SmartNIC without adding new devices. Nezha offloads stateless rule/flow tables to the remote, while keeping states locally. This eliminates the need for state synchronization, facilitating load sharing and failover. The deployment cost of Nezha is only a small fraction of that required to deploy new devices. Data collected from production show that our CPS capability bottleneck has shifted from the vSwitch to the VM kernel stack, with #concurrent flows and #vNICs increased by up to 50.4x and 40x, respectively. Xing Li 0007, Enge Song, Tian Pan 0001, Qiang Fu 0011, Yang Song 0031, Yilong Lv, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Rong Wen, Xionglie Wei, Biao Lyu, Zhigang Zong, Qinming He, Shunmin Zhu |
SIGCOMM | 15 |
| 2025 | Albatross: A Containerized Cloud Gateway Platform with FPGA-accelerated Packet-level Load BalancingabstractAlibaba Cloud's centralized gateways relied heavily on high-capacity switching ASICs, but the abrupt halt of Tofino chip evolution in Jan 2023 forced us to seek alternatives that can meet the requirements of performance, supply-chain security, code reuse, and resource efficiency. After evaluating multiple options, we developed Albatross, our 3rd gen cloud gateway based on FPGA and x86 CPUs. Albatross delivers FPGA-based packet-level load balancing to the host CPUs to prevent CPU core overload, manages large reorder buffers under high-latency jitters (100μs) during complex cloud service processing, and resolves head-of-line (HOL) blocking from packet losses or software exceptions in CPUs. To avoid being overloaded by heavy hitters due to anomalies or attacks, it also implements a two-stage rate limiter for millions of tenants with only 2MB of FPGA memory. To maximize resource utilization, Albatross uses containerization to host multiple gateway instances and designs a BGP proxy to lessen the BGP peering overhead on uplink switches caused by high-density container deployments. After hundreds of man-months of development, a single Albatross node can process 80~120Mpps of cloud network traffic with an average latency of 20μs, reducing gateway and sandbox infra costs by 50%. Jianyuan Lu, Shunmin Zhu, Tian Pan 0001, Yisong Qiao, Yang Song 0031, Wenqiang Su, Yanqiang Li, Enge Song, Shize Zhang, Xiaoqing Sun, Rong Wen, Xionglie Wei, Biao Lyu, Xing Li 0007 |
SIGCOMM | 16 |
| 2025 | ZooRoute: Enhancing Cloud-Scale Network Reliability via Overlay Proactive ReroutingabstractThis paper presents ZooRoute, a tenant-transparent, fast failure recovery service that requires no modifications to physical devices. ZooRoute leverages the overlay layer and enables traffic flows to bypass failures by altering source ports (srcPorts) in packet headers during encapsulation. To enable deployment in large-scale cloud networks, ZooRoute proposes: 1) On-demand probing to efficiently monitor a vast number of hosts while minimizing telemetry costs. 2) Table compression to record the states of numerous paths with limited on-chip resources. 3) A device-sensing mechanism to prevent unnecessary reconnections in stateful forwarding. Deployed in Alibaba Cloud for 18 months, ZooRoute has significantly improved network reliability, reducing cumulative outage time by 92.71%. Xiaoqing Sun, Xionglie Wei, Xing Li 0007, Yi Wang 0004, Chenhao Jia, Zhanlong Zhang, Jianyuan Lu, Shize Zhang, Enge Song, Yang Song 0031, Tian Pan 0001, Rong Wen, Biao Lyu, Yang Xu 0010, Shunmin Zhu |
SIGCOMM | 19 |
| 2025 | Cloud Load Balancers Need to Stay Off the Data PathabstractLoad balancers (LBs) are crucial in cloud environments, ensuring workload scalability. They route packets destined for a service (identified by a virtual IP address, or VIP) to a group of servers designated to deliver that service, each with its direct IP address (DIP). Consequently, LBs significantly impact the performance of cloud services and the experience of tenants. Many academic studies focus on specific issues such as designing new load balancing algorithms and developing hardware load balancing devices to enhance the LB's performance, reliability, and scalability. However, we believe this approach is not ideal for cloud data centers for the following reasons: (i) the increasing demands of users and the variety of cloud service types turn the LB into a bottleneck; and (ii) continually adding machines or upgrading hardware devices can incur substantial costs. In this paper, we propose the Next Generation Load Balancer (NGLB), designed to bypass the TCP connection datapath from the LB, thereby eliminating latency overheads and scalability bottlenecks of traditional cloud LBs. The LB only participates in the TCP connection establishment phase. The three key features of our design are: (i) the introduction of anactive address learningmodel to redirect traffic and bypass the LB, (ii) amulti-tenant isolationmechanism for deployment within multi-tenant Virtual Private Cloud networks, and (iii) a distributed flow control method, known ashierarchical connection cleaner, designed to ensure the availability of backend resources. The evaluation results demonstrate that NGLB reduces latency by 16% and increases nearly 3× throughput. With the same LB resources, NGLB improves 10× rate of new connection establishment. More importantly, five years of operational experience has proven NGLB's stability for high-bandwidth services. Shuai Jin, Zhenyu Wen, Shibo He, Qingzheng Hou, Yang Song 0031, Zhigang Zong, Bengbeng Xue, Ku Li, Xing Li 0007, Biao Lyu, Rong Wen, Jiming Chen 0001, Shunmin Zhu |
IEEE Trans. Cloud Comput. | 13 |
| 2025 | Enabling Stateful TCP Performance Profiling With Key Event CapturingabstractTCP ensures reliable transmission through its stateful implementation and remains crucial today. TCP performance profiling is essential for tasks like diagnosing network performance problems, optimizing transmission performance, and developing new TCP variants, etc. Existing profiling methods lack enough attention to TCP state transition to provide detailed insights on TCP performance. Thus, we build TcpSight, a tool focusing on TCP state transition throughout connection lifetimes. TcpSight conducts stateful analysis by capturing key events using an efficient per-connection lock-free data management mechanism. Besides, TcpSight enhances profiling by integrating application layer information collected from the TCP stack. With the profiling results, users can identify the culprit of TCP performance degradation, and evaluate the performance of TCP algorithms. We design optional modules and filtering mechanisms to reduce TcpSights overhead. Our evaluation presents that TcpSight incurs an additional CPU consumption of about 16.6% (without filtering) and 10.6% (with filtering) when the servers load is 55.7%, and generates storage consumption about 1.88 KB per connection on average. We also give application cases of TcpSight and the deployment experiences in Alibaba Cloud. TcpSight helps in revealing meaningful findings and insights into exploiting TCP in the production deployment. Ruopeng Geng, Jianyuan Lu, Chongrong Fang, Shaokai Zhang, Jiangu Zhao, Zhigang Zong, Biao Lyu, Shunmin Zhu, Peng Cheng 0001, Jiming Chen 0001 |
IEEE Trans. Netw. Serv. Manag. | 7 |
| 2024 | QuarkTable: Building Compact Forwarding Tables for Programmable Switches on Public CloudsabstractProgrammable switches have been recently proposed as dataplane solutions for public clouds. However, the conflict of limited on-chip memory and massive forwarding rules in cloud networks hinders the large-scale deployments. We argue that building compact forwarding tables for programmable switches is a viable option to this problem. In this paper, as a first step, we explore the feasibility of compact data structures for VPC routing tables (VRTs) and propose QuarkTable as a solution. The idea of QuarkTable builds upon the existence of redundancy in the prefixes of VRTs, supported by extensive analysis of real-world VRTs collected from six geographically distributed regions of Alibaba Cloud. By cutting the VRT into two partitions and encoding the upper part of the longer prefixes, QuarkTable effectively shortens the length of each entry and reduces the overall memory consumption. We demonstrate the effectiveness of QuarkTable by showing its proximity to the entropy bound of real-world VRTs. Experiments on six VRTs from Alibaba Cloud show practical memory savings up to 33.5% and 30.2% for SRAM and TCAM, respectively. Jianyuan Lu, Huaiyi Zhao, Yehao Feng, Shengru Li, Enge Song, Xionglie Wei, Biao Lyu, Rong Wen, Shunmin Zhu |
APNet | 9 |
| 2024 | vSwitchLB: Stratified Load Balancing for vSwitch Efficiency in Data CentersabstractThe virtual switch (vSwitch) serves as a fundamental element in cloud network, critical for high-performance and strongly isolated inter-VM forwarding in local and external networks. Similar to other multicore systems, a vSwitch with multiple cores also faces the issue of core load imbalance. As a major cloud provider, we pinpoint four cases of core load imbalance within the vSwitch in our cloud, stemming from unequal traffic distribution across virtual queues and RSS buckets, as well as from traffic patterns like heavy hitters and micro-bursts. To tackle the different load imbalance cases, we present vSwitchLB, a vSwitch load balance framework. Specifically, we introduce a load imbalance detection module, accompanied by dedicated techniques designed to address each specific type of imbalance. Our preliminary evaluation shows that vSwitchLB can accurately classify different load imbalances encountered in the vSwitch on our cloud and then prevent any single core of vSwitch from being flooded and overwhelmed. Enge Song, Yi Wang 0004, Jianyuan Lu, Xing Li 0007, Biao Lyu, Rong Wen, Shibo He, Yuanchao Shu, Shunmin Zhu |
APNet | 8 |
| 2024 | Understanding Network Startup for Secure Containers in Multi-Tenant Clouds: Performance, Bottleneck and OptimizationabstractIn this paper, we use empirical measurements to show that container network startup is a key factor that contributes to the slow startup of secure containers in multi-tenant clouds, especially in the scenario of serverless computing, where the issue is pronounced by high-volume concurrent container invocations. We conduct extensive and detailed analysis on existing Container Network Interface (CNI) plugins and show that even the fastest one doubles the startup time from the no-network scenario. We show that the major cause of the blowup in total startup time is that enabling networking significantly increases the contention among different startup stages, particularly for global Linux kernel locks, including the Routing Table NetLink (RTNL) mutex lock and various spin locks. We reveal that contending for these locks hinders startup performance in three ways, including directly increasing stage time, causing poor pipeline overlap and wasting CPU resources. To mitigate such kernel lock contention, we propose a multi-stage concurrency control mechanism based on Bayesian optimization to limit the concurrency of each contended stage. Our results show that this lightweight mechanism can effectively reduce the end-to-end container startup time by 18.8% with negligible extra overhead. Yunzhuo Liu, Junchen Guo, Bo Jiang 0003, Xiaoqing Sun, Yang Song 0031, Zhiyuan Hou, Biao Lyu, Rong Wen, Shunmin Zhu, Xinbing Wang |
IMC | 9 |
| 2024 | CloudPlanner: Minimizing Upgrade Risk of Virtual Network Devices for Large-Scale Cloud NetworksabstractCloud networks continuously upgrade softwarized virtual network devices (VNDs) to meet evolving tenant demands. However, such upgrades may result in unexpected failures. An intuitive idea to prevent upgrade failures is to resolve all compatibility issues before deployment, but it is impractical to replicate all deployed VND cases and test them with lots of replayed real traffic for the VND developers. As a result, the operations team takes upgrade risk to test upgrades by gradually deploying them. Although careful upgrade schedule planning is the most common method to minimize upgrade risk, to the best of our knowledge, no VND upgrade schedule planning scheme has been adequately studied for large-scale cloud networks. To fill this gap, we propose CloudPlanner, the first VND upgrade schedule planning scheme aiming to minimize the VND upgrade risk for large-scale cloud networks. CloudPlanner prioritizes upgrading VNDs that are more likely to trigger failures based on expert knowledge and historical failure-trigger VND properties and limits the number of tenants associated with simultaneously upgraded VNDs. We also propose a heuristic solver which can quickly and greedily plan schedules. Using real-world data from production environments, we demonstrate the benefits of CloudPlanner through extensive experiments. Enhuan Dong, Jiahai Yang 0001, Shize Zhang, Zejie Wang, Xiaoqing Sun, Enge Song, Jianyuan Lu, Biao Lyu, Shunmin Zhu |
INFOCOM | 12 |
| 2024 | LuoShen: A Hyper-Converged Programmable Gateway for Multi-Tenant Multi-Service Edge Clouds
Tian Pan 0001, Xionglie Wei, Yisong Qiao, Tiesheng Cheng, Wenqiang Su, Yuke Hong, Zhengzhong Wang, Chongjing Dai, Peiqiao Wang, Xuetao Jia, Jianyuan Lu, Enge Song, Biao Lyu, Ennan Zhai, Jiao Zhang 0002, Tao Huang 0005, Dennis Cai, Shunmin Zhu |
NSDI | 20 |
| 2024 | POSEIDON: A Consolidated Virtual Network Controller that Manages Millions of Tenants via Config Tree
Biao Lyu, Enge Song, Tian Pan 0001, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Chenxiao Wang, Xiuheng Chen, Yandong Duan, Weisheng Wang, Jinpeng Long, Kunpeng Zhou, Zhigang Zong, Xing Li 0007, Guangwang Li, Peng Cheng 0001, Jiming Chen 0001, Shunmin Zhu |
NSDI | 1 |
| 2024 | Canal Mesh: A Cloud-Scale Sidecar-Free Multi-Tenant Service Mesh ArchitectureabstractIn recent years, service mesh frameworks have gained significant popularity in building microservice-based applications. A key component of these frameworks is a proxy in each K8s pod, named sidecar, which handles inter-pod traffic. Our empirical measurement reveals that such per-pod sidecars cause numerous problems, including intrusion into the user pod, excessive resource occupation, significant overhead in managing many sidecars, and performance degradation caused by passing traffic through the sidecar. Enge Song, Yang Song 0031, Chengyun Lu, Tian Pan 0001, Shaokai Zhang, Jianyuan Lu, Jiangu Zhao, Xining Wang, Minglan Gao, Zongquan Li, Ziyang Fang, Biao Lyu, Rong Wen, Li Yi 0003, Zhigang Zong, Shunmin Zhu |
SIGCOMM | 13 |
| 2024 | CyberStar: Simple, Elastic and Cost-Effective Network Functions Management in Cloud Network at Scale
Bengbeng Xue, Yang Song 0031, Xiaoxin Peng, Yilong Lyu, Xiaoliang Wang 0001, Chen Tian 0001, Cam-Tu Nguyen, Biao Lyu, Rong Wen, Zhigang Zong, Shunmin Zhu |
USENIX ATC | 11 |
| 2024 | Performance Analysis and Power Allocation for Covert Mobile Edge Computing With RIS-Aided NOMAabstractMobile edge computing (MEC) is a key enabling technology for the sixth-generation (6G) wireless networks. In this paper, we apply covert communications to MEC to prevent information leakage, where two candidate technologies of 6G, reconfigurable intelligent surface (RIS) and non-orthogonal multiple access (NOMA), are adopted. Specifically, a legitimate transmitter sends messages to a pair of legitimate receivers, while a warden aims to detect whether the legitimate transmission exists. We can hide the existence of the stronger-signal receiver's transmission from the warden by exploiting the nature of NOMA, and we use a jammer to further hide this existence. We first analyze the performance for the case of fixed power allocation between the legitimate transmitters and the jammer. The closed-form expressions for the minimum detection error probability and ergodic public/covert rates are derived. Then, we design a reinforcement learning (RL)-based power-allocation optimization algorithm that maximizes the sum rate while ensuring covertness, by optimizing the power allocation between the transmitters and the jammer. Simulation results validate the correctness of our analysis and demonstrate the covertness of the proposed scheme. Furthermore, the performance of the RL-based algorithm is significantly better than that of the baseline scheme, which reflects the effectiveness of our proposed algorithm. Yanyu Cheng, Jianyuan Lu, Dusit Niyato, Biao Lyu, Minrui Xu, Shunmin Zhu |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | LFVeri: Network Configuration Verification for Virtual Private Cloud NetworksabstractThe Virtual Private Cloud (VPC) service enables users to configure shared resources within public clouds on demand, providing isolation between users. However, configuring the VPC network is a complex and error-prone task, and misconfiguration has been the leading cause of cloud network security issues. The large number of complex network components and configurations makes it difficult to perform scalable, efficient, and accurate fault verification of the network behavior. To address this issue, we design a comprehensive and automated fault diagnosis and localization tool, calledLFVeri, which is built upon an innovative modular network model that accurately captures the logic functions of real components within VPC networks, and propose eleven functions to verify network reachability and security requirements. We conduct performance testing ofLFVerion various datasets and compared it with other verification tools. The experiments show thatLFVerioutperforms in modeling and analyzing real VPC scenarios while also possessing the fastest verification speed. It can model and analyze large VPC networks with tens of thousands of components and millions of configuration rules in less than half an hour. Kun Wang 0023, Chengcheng Zhao, Jinpei Chu, Yiping Shi, Jianyuan Lu, Biao Lyu, Shunmin Zhu, Peng Cheng 0001, Jiming Chen 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2024 | Proactive Telemetry in Large-Scale Multi-Tenant Cloud Overlay NetworksabstractAt present, public clouds have served millions of tenants. To provide reliable services, cloud vendors need to perceive health status of the cloud network by building a telemetry system to detect possible network failures. While telemetry systems for physical networks have been extensively studied, research on telemetry systems for virtual networks is still insufficient. Different from physical networks, we conclude that building a virtual network telemetry system faces new challenges of feasibility, efficiency, and effectiveness. Specifically, we need to 1) protect privacy of tenants and adapt to heterogeneous middleboxes at the data plane; 2) handle frequent virtual network topology updates and compress large-scale measurement paths for millions of tenants at the control plane; 3) analyze telemetry results to locate network failures at the analysis plane. To address these challenges, we present Zoonet, a proactive virtual network telemetry system for multi-tenant clouds. At the data plane, Zoonet uses host agent and arp-ping to protect tenants’ privacy and defines an elegant generalization of ping and traceroute, which can work on heterogeneous middleboxes. At the control plane, Zoonet conducts update batch processing and substantial probing path pruning to lessen the overhead. At the analysis plane, Zoonet reduces noises and aggregates alerts based on temporal and spatial correlation and conducts the hop-by-hop telemetry mode to locate failures. Zoonet has been deployed in Alibaba Cloud for over two years, covering tens of cloud regions, hundreds of thousands of servers. We become increasingly reliant on Zoonet as it reduces 86% of the personnel engaged in troubleshooting. Shunmin Zhu, Jianyuan Lu, Biao Lyu, Tian Pan 0001, Shize Zhang, Xiaoqing Sun, Chenhao Jia, Xin Cheng 0022, Daxiang Kang, Yilong Lv, Fukun Yang, Xiaobo Xue, Xihui Yang, Jiahai Yang 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2024 | CloudSentry: Two-Stage Heavy Hitter Detection for Cloud-Scale Gateway Overload ProtectionabstractThe cloud vendors provide sharing resources for millions of tenants across the world to achieve economies of scale. At the same time, the cloud network keeps the performance isolation between different tenants as if they use their private dedicated resources. However, heavy hitters caused by a single tenant at cloud gateways will break such isolation, undermining the predictable performance expected by other cloud tenants. To prevent it, heavy hitter detection becomes a key concern at the performance-critical cloud gateways but faces the dilemma between fine granularity and low overhead. In this work, we presentCloudSentry, a scalable two-stage heavy hitter detection system dedicated to multi-tenant cloud gateways against such a dilemma. CloudSentry uses CPU utilization as an indicator of heavy hitters and conducts a lightweight coarse-grained detection running 24/7 to detect such CPU spikes. Then it invokes a fine-grained detection to precisely dump and analyze the potential heavy-hitter packets at the CPU spikes. After that, a more comprehensive analysis is conducted to associate heavy hitters with the cloud service scenarios and invoke a corresponding backpressure procedure. CloudSentry significantly reduces memory, computation and storage overhead compared with existing approaches. In a gateway cluster under an average traffic throughput of 251 Gbps, CloudSentry consumes only a fraction of 2%–5% CPU utilization with 8 KB run-time memory, producing only 10 MB heavy hitter logs during one month. Additionally, as it has been deployed in Alibaba Cloud for over two years, we share case studies and a lot of deployment experiences in this article. Jianyuan Lu, Tian Pan 0001, Mao Miao, Guangzhe Zhou, Yining Qi, Shize Zhang, Enge Song, Xiaoqing Sun, Huaiyi Zhao, Biao Lyu, Shunmin Zhu |
IEEE Trans. Parallel Distributed Syst. | 11 |
| 2024 | CouldPin-Fast: Effient and Effective Root Cause Localization for Shared Bandwidth Package Traffic Anomalies in Public Cloud NetworksabstractAs cloud services become increasingly widespread, many public cloud tenants opt for Shared Bandwidth Package (sBwp) services for inbound/outbound communication. The sBwp service allows tenants to purchase shared bandwidth for multiple virtual machines (VMs) instead of buying it individually, which is a convenient and cost-effective traffic management mode. However, the sBwp service presents new challenges for operators to identify the root cause of abnormal sBwp traffic, especially in large-scale, globally distributed public clouds with millions of users. Developing a localization system in public cloud faces several challenges, including dynamic scalability, hyper-scale data efficiently obtaining, and complex application scenarios. To address these challenges, we propose a two-stage localization method calledCloudPin-Fast. First,CloudPin-Fastemploys a cold-start mode to meet dynamic requirements. Second,CloudPin-Fastimplements a pre-filter to reduce the transmission and processing of hyper-scale data. Finally,CloudPin-Fastuses an anomaly localization algorithm based on multi-dimensional statistics fusion in the second stage to cover complex scenarios. The evaluation results on four production datasets have shown superior efficiency and effectiveness. We also share lessons learned from deployingCloudPin-Fastfor over a year in a world-renowned public cloud vendor. Shize Zhang, Jianyuan Lu, Biao Lyu, Shunmin Zhu, Enhuan Dong, Jiahai Yang 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2023 | X-Plane: A High-Throughput Large-Capacity 5G UPFabstractCloud providers, such as AWS and Azure, have started providing 5G services on their cloud infrastructure. In this paper, we present the design and implementation of X-Plane, a system that uses commercial programmable ASICs and DRAM servers on today's cloud infrastructure to implement high-performance 5G User Plane Function (UPF). Building X-Plane is hard because we need to address the following challenges: consistency issues when concurrently accessing UPF state data, slow UE table lookup due to repetitive and numerous Packet Detection Rule (PDR) matching, and the need to handle out-of-order packets from disconnected UEs. X-Plane addresses these challenges by designing three novel technologies: concurrent state data access protocol, fast flow table and paging buffer for handling out-of-order packets. We demonstrate its feasibility and practicality with our implementation on a Tofino-based programmable ASIC. Our evaluation shows that X-Plane can support over ~490Gbps throughput per ASIC pipeline, over 10 million UEs, and finish packet processing within predictable ~4 us on average. Yunzhuo Liu, Hao Nie, Bo Jiang 0003, Yirui Liu 0001, Yidong Yao, Xionglie Wei, Biao Lyu, Chenren Xu, Shunmin Zhu, Xinbing Wang |
MobiCom | 9 |
| 2023 | FlowPinpoint: Localizing Anomalies in Cloud-Client Services for Cloud ProvidersabstractFor public cloud providers, it is of great significance to maintain the availability of their cloud services, which requires efficient anomaly diagnosis and recovery. To achieve such properties, the first step is to localize the anomalies, i.e., determining where they happen in the network path of cloud-client services. We propose FlowPinpoint to perform anomaly localization for cloud providers. FlowPinpoint collects statistics of each network flow at the cloud network gateways (i.e., gateway flowlog), where the collected data can reflect the information from both the cloud side and the Internet side. Aggregation and association are conducted on the datacenter-scale gateway flowlogs by Alibaba's big data computing platform. In order to preclude the disturbance of anomaly-unrelated flowlogs, a two-layer filter is proposed which consists of an indicator-based filter and an isolation forest filter. Finally, the anomaly localization analyzer classifies the flowlogs and determines whether the anomaly is inside the cloud network or not according to the classification results. FlowPinpoint is implemented and tested in the production environment of Alibaba Cloud, and it correctly localizes 1 anomaly inside the cloud and 6 anomalies on the Internet over 4 months. Ruopeng Geng, Chongrong Fang, Shiyang Guo, Daxiang Kang, Biao Lyu, Shunmin Zhu, Peng Cheng 0001 |
IEEE Trans. Cloud Comput. | 5 |
| 2022 | Zoonet: a proactive telemetry system for large-scale cloud networksabstractWe present Zoonet, a proactive virtual network telemetry system for multi-tenant clouds. The requirements are to (1) cover hyper-scale virtual networks with millions of tenants and millions of VMs for top tenants; (2) handle frequent virtual topology changes due to tenants' configuration through flexible APIs; (3) adapt to heterogeneous middleboxes along the probing paths; (4) achieve VM-to-VM telemetry without breaking tenant privacy; (5) differentiate virtual and physical network problems. We argue existing physical network telemetry solutions fail to satisfy our needs due to either incomplete telemetry coverage or outrageous telemetry overhead. Zoonet sets an ambitious goal to provide VM-to-VM hop-by-hop telemetry for each tenant, which is achieved based on self-developed, customizable middleboxes via hundreds of person-months under close team collaboration. At the data plane, Zoonet defines an elegant generalization of ping and traceroute, but made to work on multi-tenant clouds with heterogeneous middleboxes. At the control plane, Zoonet conducts substantial probing path pruning and update batch processing to lessen the overhead. Zoonet has been deployed in Alibaba Cloud for over two years, covering tens of cloud regions, hundreds of thousands of servers. We become increasingly reliant on Zoonet as it reduces 86% of the personnel engaged in troubleshooting. Shunmin Zhu, Jianyuan Lu, Biao Lyu, Tian Pan 0001, Chenhao Jia, Xin Cheng 0022, Daxiang Kang, Yilong Lv, Fukun Yang, Xiaobo Xue, Jiahai Yang 0001 |
CoNEXT | 3 |
| 2022 | Performance Analysis of Jammer-Aided Covert RIS-NOMA SystemsabstractIn this paper, we apply covert communications to reconfigurable intelligent surface (RIS)-assisted non-orthogonal multiple access (NOMA) networks, where a legitimate transmitter sends messages to a pair of legitimate users while a warden aims to detect whether the legitimate transmission exists. We can hide the existence of the strong user's transmission from a warden by exploiting the nature of NOMA, i.e., allocating less power to the strong user, and we use a jammer to further hide that existence. Correspondingly, we analyze the system performance and obtain the closed-form expression for the minimum detection error probability. Simulation results validate the correctness of our analysis and demonstrate the covertness of the proposed scheme. Yanyu Cheng, Jianyuan Lu, Dusit Niyato, Biao Lyu, Minrui Xu, Shunmin Zhu |
GLOBECOM | 4 |
| 2022 | MIMIC: SmartNIC-aided Flow Backpressure for CPU Overloading Protection in Multi-Tenant CloudsabstractIn multi-tenant clouds, off-the-shelf x86 boxes are widely deployed as middleboxes. With the rapid growth of cloud traffic and the migration to NFV deployment in recent years, CPU overloading at middleboxes becomes more of an issue. From our data centers, we observed that the CPU overloading was caused by heavy hitters. To address this issue, we propose MIMIC, a cloud-scale flow backpressure system, implemented onto our existing SmartNIC with FPGA acceleration. MIMIC rate-limits the selected heavy hitters through a new per-flow backpressure protocol and a new heavy-hitter detection system, to protect the other tenants. The detection system is based on hierarchical memory design, leveraging on-chip SRAM and off-chip DRAM, which can handle highly concurrent cloud traffic without the losses of flow information. We extend the design by adding a pre-filtering procedure for rapid detection. To avoid CPU being flooded by FPGA through frequent heavy-hitter reporting, due to their performance disparity, the CPU queries the FPGA on demand. The backpressure protocol is non-invasive to protect tenant privacy and allows controllable rate-limiting through the novel use of ECN and meter tables. The SmartNIC acts as a man in the middle to facilitate heavy-hitter detection and per-flow backpressuring. In a production setting, we observe that MIMIC can react quickly and bring down CPU load to the normal level within 10ms without packet losses. Enge Song, Nianbing Yu, Tian Pan 0001, Qiang Fu 0011, Xionglie Wei, Yisong Qiao, Jianyuan Lu, Yijian Dong, Mingxu Xie, Jinkui Mao, Zhengjie Luo, Chenhao Jia, Jiao Zhang 0002, Tao Huang 0005, Biao Lyu, Shunmin Zhu |
ICNP | 17 |
| 2022 | Towards Automatic Root Cause Diagnosis of Persistent Packet Loss in Cloud Overlay NetworkabstractPersistent packet loss in the cloud-scale overlay network severely compromises tenant experiences. Cloud providers are keen to diagnose such problems efficiently. However, existing work is either designed for the physical network or insufficient to present the concrete reason of packet loss. We propose to record and analyze the on-site forwarding condition of packets during packet-level tracing. The cloud-scale overlay network presents great challenges to achieve this goal with its high network complexity, multi-tenant nature, and diversity of root causes. To address these challenges, we present VTrace, an automatic diagnostic system for persistent packet loss over the cloud-scale overlay network. Utilizing the “fast path-slow path” structure of virtual forwarding devices (VFDs), e.g., vSwitches, VTrace installs several “coloring-matching-logging” rules in VFDs to selectively track the target packets and inspect them in depth. The detailed forwarding situation at each hop is logged and then assembled to perform analysis with an efficient path reconstruction scheme. Experiments are conducted to demonstrate VTrace’s low overhead and quick response. Besides, based on the idea “coloring-matching-counting”, VTrace can be easily extended toVTrace-statsto identify the culprit device for transient packet loss. We share experiences of how VTrace andVTrace-statsefficiently work after deploying them in Alibaba Cloud for years. Chongrong Fang, Haoyu Liu 0002, Mao Miao, Lei Wang 0005, Wansheng Zhang, Daxiang Kang, Biao Lyu, Shunmin Zhu, Peng Cheng 0001, Jiming Chen 0001 |
IEEE/ACM Trans. Netw. | 8 |
| 2021 | A Two-Stage Heavy Hitter Detection System Based on CPU Spikes at Cloud-Scale GatewaysabstractThe cloud network provides sharing resources for tens of thousands of tenants to achieve economics of scale. However, heavy hitters caused by a single tenant will probably interfere with the processing of the cloud gateways, undermining the predictable performance expected by other cloud tenants. To prevent it, heavy hitter detection becomes a key concern at the performance-critical cloud gateways but faces the dilemma between fine granularity and low overhead. In this work, we present CloudSentry, a scalable two-stage heavy hitter detection system dedicated to multi-tenant cloud gateways against such a dilemma. CloudSentry contains a lightweight coarse-grained detection running 24/7 to localize infrequent CPU spikes. Then it invokes a fine-grained detection to precisely dump and analyze the potential heavy-hitter packets at the CPU spikes. After that, a more comprehensive analysis is conducted to associate heavy hitters with the cloud service scenarios and invoke a corresponding backpressure procedure. CloudSentry significantly reduces memory, computation and storage overhead compared with existing approaches. Additionally, it has been deployed world-wide in Alibaba Cloud for over one year, with rich deployment experiences. In a gateway cluster under an average traffic throughput of of 251Gbps, CloudSentry consumes only a fraction of 2%-5% CPU utilization with 8KB run-time memory, producing only 10MB heavy hitter logs during one month. Jianyuan Lu, Tian Pan 0001, Mao Miao, Guangzhe Zhou, Yining Qi, Biao Lyu, Shunmin Zhu |
ICDCS | 7 |
| 2021 | CloudPin: A Root Cause Localization Framework of Shared Bandwidth Package Traffic Anomalies in Public Cloud NetworksabstractDue to the sharing nature of public cloud, most of the cloud services use a sharing bandwidth package (sBwp) model to conduct inbound/outbound communication. The sBwp model allows users to purchase a sharing bandwidth for plenty of virtual machines instead of purchasing bandwidth for each virtual machine separately. The advantage of sBwp is that it can provide users with convenient configuration and lower economic cost. However, the sBwp model brings new challenges for operators to localize the root cause of traffic anomalies of a sharing bandwidth, especially for a globally distributed large-scale public cloud with millions of users. In this paper, we first formalize the sBwp problem on the cloud and propose CloudPin, a root cause localization framework for this problem. Our framework solves all the challenges by employing a multi-dimensional algorithm with three sub-models of prediction deviation, anomaly ampli-tude, and shape similarity, and an overall ranking algorithm. Evaluations on real-world data, from one of the world-renowned public cloud vendors, show that our algorithm precision reaches 97.8% for the top 1 of the ranking list, outperforming multiple baseline algorithms. Shize Zhang, Jianyuan Lu, Biao Lyu, Shunmin Zhu, Jiahai Yang 0001, Lin He 0004 |
ISSRE | 4 |
| 2021 | A survey of cloud network fault diagnostic systems and toolsabstractRecently, cloud computing has become a vital part that supports people’s normal lives and production. However, accompanied by the increasing complexity of the cloud network, failures constantly keep coming up and cause huge economic losses. Thus, to guarantee the cloud network performance and prevent execrable effects caused by failures, cloud network diagnostics has become of great interest for cloud service providers. Due to the characteristics of cloud network (e.g., virtualization and multi-tenancy), transplanting traditional network diagnostic tools to the cloud network face several difficulties. Additionally, many existing tools cannot solve problems in the cloud network. In this paper, we summarize and classify the state-of-the-art technologies of cloud diagnostics which can be used in the production cloud network according to their features. Moreover, we analyze the differences between cloud network diagnostics and traditional network diagnostics based on the characteristics of the cloud network. Considering the operation requirements of the cloud network, we propose the points that should be cared about when designing a cloud network diagnostic tool. Also, we discuss the challenges that cloud network diagnostics will face in future development. Yining Qi, Chongrong Fang, Haoyu Liu 0002, Daxiang Kang, Biao Lyu, Peng Cheng 0001, Jiming Chen 0001 |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2020 | RAIN: Towards Real-Time Core Devices Anomaly Detection Through Session Data in Cloud NetworkabstractCore devices form the critical components of the cloud network and provide service to multiple tenants simultaneously. The anomalies that happened in core devices impact network availability of a large number of users, meanwhile, lead to the degradation of cloud providers’ profits. However, direct monitoring of core devices needs to deploy massive heartbeat checking tools on numerous related components, which will be extremely laborious. In this paper, we deploy RAIN to reduce the number of devices that need to be detailed investigated for anomalies. The session traffic data among core devices and served virtual machines are utilized to conduct the analyzing. To guarantee near real-time monitoring, RAIN is designed as a two-step structure and incorporating four feature-based detection methods. RAIN has been deployed in Alibaba’s production cloud network for over 6 months and is analyzing terabytes of traffic flow metrics per day. Haoyu Liu 0002, Chongrong Fang, Yining Qi, Shaozhe Wang, Daxiang Kang, Biao Lyu, Peng Cheng 0001, Jiming Chen 0001 |
NOMS | 8 |
| 2020 | VTrace: Automatic Diagnostic System for Persistent Packet Loss in Cloud-Scale Overlay NetworkabstractPersistent packet loss in the cloud-scale overlay network severely compromises tenant experiences. Cloud providers are keen to automatically and quickly determine the root cause of such problems. However, existing work is either designed for the physical network or insufficient to present the concrete reason of packet loss. In this paper, we propose to record and analyze the on-site forwarding condition of packets during packet-level tracing. The cloud-scale overlay network presents great challenges to achieve this goal with its high network complexity, multi-tenant nature, and diversity of root causes. To address these challenges, we present VTrace, an automatic diagnostic system for persistent packet loss over the cloud-scale overlay network. Utilizing the "fast path-slow path" structure of virtual forwarding devices (VFDs), e.g., vSwitches, VTrace installs several "coloring, matching and logging" rules in VFDs to selectively track the packets of interest and inspect them in depth. The detailed forwarding situation at each hop is logged and then assembled to perform analysis with an efficient path reconstruction scheme. Experiments are conducted to demonstrate VTrace's low overhead and quick responsiveness. We share experiences of how VTrace efficiently resolves persistent packet loss issues after deploying it in Alibaba Cloud for over 20 months. Chongrong Fang, Haoyu Liu 0002, Mao Miao, Lei Wang 0005, Wansheng Zhang, Daxiang Kang, Biao Lyu, Peng Cheng 0001, Jiming Chen 0001 |
SIGCOMM | 8 |