Yang Song 0031

dblp:24/4470-31 · DBLP profile ↗
← Back
21ranked-venue papers
0as first author
20since 2021 · last 2026
0009-0008-2611-757XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 18 · 17 since 2021Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Scaling LLM Agent Tool Access at Cloud Scale
abstract
LLM agents increasingly rely on tool calling, and the Model Context Protocol (MCP) standardizes it between agents and tool providers, reducing integration cost and driving rapid growth in tool scale. Yet a standardized interface does not make tool access work at production scale: legacy services are not MCP-callable, fast protocol evolution creates compatibility cost, large tool sets exhaust the context window, and stateful sessions complicate load balancing. We solve these with a shared control point, a centralized MCP Gateway System that makes MCP operational at cloud scale. The gateway breaks the direct-connect data plane and consolidates legacy API integration, protocol bridging, access control, and session-aware routing, while scaling out elastically at low per-call overhead. It scales agent tool access to thousands of cloud operations.
Enge Song, Yueshang Zuo, Rong Wen, Jing Tie, Zhou Shao, Qiang Fu 0011, Xiaobo Xue, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Tian Pan 0001, Yang Song 0031, Xing Li 0007, Biao Lyu, Meng Li 0010, Haipeng Dai 0001, Guihai Chen, Shunmin Zhu
APNet23
2026 Integrating AI Clusters into Virtual Private Cloud
abstract
While commodity NIC-based back-end AI networks offer ultra-high intra-cluster bandwidth for distributed training, their limited programmability and on-chip resources hinder the implementation of advanced VPC features such as fine-grained isolation and stateful security policies. Furthermore, access to resources within the VPC needs to be routed through the front-end DPU, which is shared by the scale-up domain. The mismatch between the front-end DPU’s bandwidth and the back-end requirements causes GPU underutilization when intensive VPC communication is required for content recommendation, AIGC, and federated learning workloads. We propose an architecture that decouples complex policy enforcement from high-speed packet forwarding to support VPC semantics on back-end NICs and enable front-end/back-end integration. Evaluations show near-full GPU utilization in our analytical model and 71 μ s P999 extra latency of the first packet, suggesting that commodity hardware can support both high-throughput AI training and flexible VPC features.
Xing Li 0007, Enge Song, Changgang Zheng, Shengyao Gao, Juncheng Xiang, Junnan Cai, Haoxiang Pan, Yang Song 0031, Yilong Lv, Qiang Fu 0011, Zhigang Zong, Shunmin Zhu
APNet13
2026 Zephyr: A Zero-loss and Tranparent TLS Connection Migration Framework
abstract
While essential for stateful modern workloads like Large Language Model agents and IoT services, long-lived connections impede cloud infrastructure agility by complicating maintenance and load balancing. Existing connection migration solutions either lack support for industrial-grade encrypted traffic or fail to prevent packet loss during handover in active production environments. To address this gap, we propose Zephyr, a zero-loss and transparent TLS connection migration framework for cross-node migration between servers with different addresses. Zephyr ensures transport-layer consistency by orchestrating an eBPF-based packet buffering mechanism to safely intercept in-flight data. At the application layer, rather than deeply modifying standard TLS libraries, Zephyr creatively reuses the native session resumption mechanism via a “fake client” strategy to reconstruct complex cryptographic states without client involvement. Implemented in widely-used industrial stacks (Nginx and OpenSSL), Zephyr achieves connection migration with approximately 4.1 ms downtime and strict zero packet loss. This approach enables seamless infrastructure optimization without disrupting cloud services.
Chengcheng Yu, Yueshang Zuo, Enge Song, Shaokai Zhang, Jiangu Zhao, Tian Pan 0001, Yang Song 0031, Xing Li 0007, Rong Wen, Chengkun Wei, Shunmin Zhu, Wenzhi Chen
APNet8
2026 Single-Core Hotspots on Your VNF? Break Them Up!
abstract
Current NFVs assign packets to CPU cores at flow granularity, where each flow is pinned to a single CPU. This approach is efficient under most scenarios but has exposed limitations when handling elephant flows. These “heavy hitters” overwhelm single cores, creating bottlenecks that affect overall throughput and degrade service quality. As networks scale to higher-speed links and core-rich CPUs, these imbalances become more severe. In this paper, we propose ParaFlowO, an architecture that Parallelizes processing elephant Flows across multiple CPU cores while preserving in-Order delivery. ParaFlowO breaks elephant flows into flowlets and dynamically rotates them across multiple cores. It integrates a lightweight reordering mechanism to preserve packet order and controls parallelism to mitigate contention on shared state. Preliminary evaluations show that ParaFlowO offers a practical solution to mixed-grained parallelism in stateful middleboxes.
Changgang Zheng, Jin Ke 0005, Enge Song, Yilong Lv, Yisong Qiao, Donglin Lai, Bengbeng Xue, Yang Song 0031, Xing Li 0007, Rong Wen, Zhigang Zong, Shunmin Zhu
APNet14
2026 SOPSmith: Forging Executable SOPs for LLM-Driven GPU Cluster Network Diagnosis
Guoyao Yu, Xiaoqing Sun, Yangyang Shi, Yang Song 0031, Xing Li 0007, Biao Lyu, Zhenguang Liu, Qinming He
IWQoS6
2026 Bifrost: Alibaba's Next-Generation VPC Network with High-Performance Multipath Reliable Transport
Xing Li 0007, Bo Jiang 0003, Yilong Lv, Yuke Hong, Yinian Zhou, Junnan Cai, Jiayue Xu, Yunrui Hu, Zhao Gao, Enge Song, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Changgang Zheng, Yang Song 0031, Biao Lyu, Rong Wen, Zhigang Zong, Shunmin Zhu
NSDI23
2026 CStar Gateway: Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Tian Pan 0001, Jin Ke 0005, Baohai Hu, Changgang Zheng, Enge Song, Donglin Lai, Yisong Qiao, Bengbeng Xue, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Yang Song 0031, Xionglie Wei, Biao Lyu, Rong Wen, Zhigang Zong, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu
NSDI18
2026 ZooRoute: Enhancing Cloud-Scale Network Reliability via Candidate Path Provisioning and Overlay Proactive Rerouting
Xiaoqing Sun, Xing Li 0007, Xionglie Wei, Tian Pan 0001, Yi Wang 0004, Chenhao Jia, Zhanlong Zhang, Xiaobo Xue, Jianyuan Lu, Shize Zhang, Enge Song, Yang Song 0031, Rong Wen, Biao Lyu, Yang Xu 0010, Shunmin Zhu
NSDI18
2026 Spillway: Orchestrating DPU and Host into a Unified vSwitching Fabric
abstract
The transition to Data Processing Unit (DPU)-centric architectures has become the de-facto standard in modern cloud networks, enabling infrastructure offload and improved host resource utilization. However, the fixed hardware limits of DPUs increasingly fail to keep pace with the rapid growth of host compute density and network-intensive workloads. As a result, when DPU resources are saturated, host compute capacity often remains underutilized due to insufficient network provisioning.
Xiaochong Jiang, Yilong Lv, Naixuan Guan, Qiming Zhao, Sihan Fu, Xuyang Ge, Denghui Wu, Yibin Shen, Guochun Hong, Yijian Dong, Yiquan Chen, Shaoliang An, Zhixiong Guo, Yisong Qiao, Hongwei Ding 0004, Shize Zhang, Rong Wen, Yang Song 0031, Zhigang Zong, Xing Li 0007, Chengkun Wei, Shunmin Zhu, Wenzhi Chen
SIGCOMM30
2025 Understanding the Long Tail Latency of TCP in Large-Scale Cloud Networks
Enge Song, Bo Jiang 0003, Yang Song 0031, Yuke Hong, Yilong Lv, Yinian Zhou, Junnan Cai, Chao Wang 0128, Yi Wang 0004, Yehao Feng, Shize Zhang, Xiaoqing Sun, Jianyuan Lu, Xing Li 0007, Biao Lyu, Zhigang Zong, Shunmin Zhu
APNet4
2025 Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Yang Song 0031, Tian Pan 0001, Zhigang Zong, Bengbeng Xue, Xionglie Wei, Yisong Qiao, Donglin Lai, Baohai Hu, Jin Ke 0005, Enge Song, Jianyuan Lu, Xing Li 0007, Biao Lyu, Rong Wen, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu
APNet2
2025 FastIOV: Fast Startup of Passthrough Network I/O Virtualization for Secure Containers
abstract
Single Root I/O Virtualization (SR-IOV) technology has advanced in recent years and can simultaneously satisfy the network requirements of high data plane performance, high deployment density, and fast startup for applications in traditional containers. However, it falls short with secure containers, which have become the mainstream choice in multi-tenant clouds. SR-IOV requires secure containers to use passthrough I/O for higher data plane performance, which hinders the container startup performance and prevents its usage in time-sensitive tasks like serverless computing. In this paper, we advocate that the startup performance of SR-IOV enabled secure containers can be further boosted, making SR-IOV suitable for building a Container Network Interface (CNI) for secure containers. We first dissect the end-to-end concurrent startup process and identify three key bottlenecks that lead to the slow startup, including Virtual Function I/O device set management, Direct Memory Access memory mapping, and Virtual Function (VF) driver initialization. We then propose a CNI named FastIOV that addresses these bottlenecks through lock decomposition, unnecessary mapping skipping, decoupled zeroing, and asynchronous VF driver initialization. Our evaluation shows that FastIOV reduces the overhead of enabling SR-IOV for secure containers by 96.1%, achieving 65.7% and 75.4% reductions in the average and 99th percentile end-to-end startup time.
Yunzhuo Liu, Junchen Guo, Bo Jiang 0003, Yang Song 0031, Rong Wen, Biao Lyu, Shunmin Zhu, Xinbing Wang
EuroSys4
2025 Hermes: Enhancing Layer-7 Cloud Load Balancers with Userspace-Directed I/O Event Notification
abstract
Layer-7 load balancers (L7 LBs) improve service performance, availability, and scalability in public clouds. They rely on I/O event notification mechanisms such as epoll to dispatch connections from the kernel to userspace workers. However, early epoll versions suffered from the thundering herd problem. Epoll exclusive (available since Linux 4.5) mitigates this but introduces LIFO wakeups, causing connection concentration on a few workers. Reuseport (Linux 3.9) hashes connections across workers but suffers from hash collisions and lacks awareness of worker load. Since each worker serves multi-tenant traffic, inter-worker load balancing is critical to avoid worker overload and preserve tenant performance isolation.
Tian Pan 0001, Enge Song, Yueshang Zuo, Shaokai Zhang, Yang Song 0031, Jiangu Zhao, Wengang Hou, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Jiao Zhang 0002, Tao Huang 0005, Biao Lyu, Xing Li 0007, Rong Wen, Zhigang Zong, Shunmin Zhu
SIGCOMM5
2025 Nezha: SmartNIC-based Virtual Switch Load Sharing
abstract
Cloud providers use SmartNIC-accelerated virtual switches (vSwitches) to offer rich network functions (NFs) for tenant VMs. Constrained by limited SmartNIC resources, it is a challenge to provide sufficient network performance for high-demand VMs. Meanwhile, we observed a significant number of idle vSwitches in the data center, which led us to consider leveraging them to build a remote resource pool for high-demand virtual NICs (vNICs). In this work, we propose Nezha, a distributed vSwitch load sharing system. Nezha reuses the existing idle SmartNICs to handle the excess load from the local SmartNIC without adding new devices. Nezha offloads stateless rule/flow tables to the remote, while keeping states locally. This eliminates the need for state synchronization, facilitating load sharing and failover. The deployment cost of Nezha is only a small fraction of that required to deploy new devices. Data collected from production show that our CPS capability bottleneck has shifted from the vSwitch to the VM kernel stack, with #concurrent flows and #vNICs increased by up to 50.4x and 40x, respectively.
Xing Li 0007, Enge Song, Tian Pan 0001, Qiang Fu 0011, Yang Song 0031, Yilong Lv, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Rong Wen, Xionglie Wei, Biao Lyu, Zhigang Zong, Qinming He, Shunmin Zhu
SIGCOMM7
2025 Albatross: A Containerized Cloud Gateway Platform with FPGA-accelerated Packet-level Load Balancing
abstract
Alibaba Cloud's centralized gateways relied heavily on high-capacity switching ASICs, but the abrupt halt of Tofino chip evolution in Jan 2023 forced us to seek alternatives that can meet the requirements of performance, supply-chain security, code reuse, and resource efficiency. After evaluating multiple options, we developed Albatross, our 3rd gen cloud gateway based on FPGA and x86 CPUs. Albatross delivers FPGA-based packet-level load balancing to the host CPUs to prevent CPU core overload, manages large reorder buffers under high-latency jitters (100μs) during complex cloud service processing, and resolves head-of-line (HOL) blocking from packet losses or software exceptions in CPUs. To avoid being overloaded by heavy hitters due to anomalies or attacks, it also implements a two-stage rate limiter for millions of tenants with only 2MB of FPGA memory. To maximize resource utilization, Albatross uses containerization to host multiple gateway instances and designs a BGP proxy to lessen the BGP peering overhead on uplink switches caused by high-density container deployments. After hundreds of man-months of development, a single Albatross node can process 80~120Mpps of cloud network traffic with an average latency of 20μs, reducing gateway and sandbox infra costs by 50%.
Jianyuan Lu, Shunmin Zhu, Tian Pan 0001, Yisong Qiao, Yang Song 0031, Wenqiang Su, Yanqiang Li, Enge Song, Shize Zhang, Xiaoqing Sun, Rong Wen, Xionglie Wei, Biao Lyu, Xing Li 0007
SIGCOMM7
2025 ZooRoute: Enhancing Cloud-Scale Network Reliability via Overlay Proactive Rerouting
abstract
This paper presents ZooRoute, a tenant-transparent, fast failure recovery service that requires no modifications to physical devices. ZooRoute leverages the overlay layer and enables traffic flows to bypass failures by altering source ports (srcPorts) in packet headers during encapsulation. To enable deployment in large-scale cloud networks, ZooRoute proposes: 1) On-demand probing to efficiently monitor a vast number of hosts while minimizing telemetry costs. 2) Table compression to record the states of numerous paths with limited on-chip resources. 3) A device-sensing mechanism to prevent unnecessary reconnections in stateful forwarding. Deployed in Alibaba Cloud for 18 months, ZooRoute has significantly improved network reliability, reducing cumulative outage time by 92.71%.
Xiaoqing Sun, Xionglie Wei, Xing Li 0007, Yi Wang 0004, Chenhao Jia, Zhanlong Zhang, Jianyuan Lu, Shize Zhang, Enge Song, Yang Song 0031, Tian Pan 0001, Rong Wen, Biao Lyu, Yang Xu 0010, Shunmin Zhu
SIGCOMM16
2025 Cloud Load Balancers Need to Stay Off the Data Path
abstract
Load balancers (LBs) are crucial in cloud environments, ensuring workload scalability. They route packets destined for a service (identified by a virtual IP address, or VIP) to a group of servers designated to deliver that service, each with its direct IP address (DIP). Consequently, LBs significantly impact the performance of cloud services and the experience of tenants. Many academic studies focus on specific issues such as designing new load balancing algorithms and developing hardware load balancing devices to enhance the LB's performance, reliability, and scalability. However, we believe this approach is not ideal for cloud data centers for the following reasons: (i) the increasing demands of users and the variety of cloud service types turn the LB into a bottleneck; and (ii) continually adding machines or upgrading hardware devices can incur substantial costs. In this paper, we propose the Next Generation Load Balancer (NGLB), designed to bypass the TCP connection datapath from the LB, thereby eliminating latency overheads and scalability bottlenecks of traditional cloud LBs. The LB only participates in the TCP connection establishment phase. The three key features of our design are: (i) the introduction of anactive address learningmodel to redirect traffic and bypass the LB, (ii) amulti-tenant isolationmechanism for deployment within multi-tenant Virtual Private Cloud networks, and (iii) a distributed flow control method, known ashierarchical connection cleaner, designed to ensure the availability of backend resources. The evaluation results demonstrate that NGLB reduces latency by 16% and increases nearly 3× throughput. With the same LB resources, NGLB improves 10× rate of new connection establishment. More importantly, five years of operational experience has proven NGLB's stability for high-bandwidth services.
Shuai Jin, Zhenyu Wen, Shibo He, Qingzheng Hou, Yang Song 0031, Zhigang Zong, Bengbeng Xue, Ku Li, Xing Li 0007, Biao Lyu, Rong Wen, Jiming Chen 0001, Shunmin Zhu
IEEE Trans. Cloud Comput.6
2024 Understanding Network Startup for Secure Containers in Multi-Tenant Clouds: Performance, Bottleneck and Optimization
abstract
In this paper, we use empirical measurements to show that container network startup is a key factor that contributes to the slow startup of secure containers in multi-tenant clouds, especially in the scenario of serverless computing, where the issue is pronounced by high-volume concurrent container invocations. We conduct extensive and detailed analysis on existing Container Network Interface (CNI) plugins and show that even the fastest one doubles the startup time from the no-network scenario. We show that the major cause of the blowup in total startup time is that enabling networking significantly increases the contention among different startup stages, particularly for global Linux kernel locks, including the Routing Table NetLink (RTNL) mutex lock and various spin locks. We reveal that contending for these locks hinders startup performance in three ways, including directly increasing stage time, causing poor pipeline overlap and wasting CPU resources. To mitigate such kernel lock contention, we propose a multi-stage concurrency control mechanism based on Bayesian optimization to limit the concurrency of each contended stage. Our results show that this lightweight mechanism can effectively reduce the end-to-end container startup time by 18.8% with negligible extra overhead.
Yunzhuo Liu, Junchen Guo, Bo Jiang 0003, Xiaoqing Sun, Yang Song 0031, Zhiyuan Hou, Biao Lyu, Rong Wen, Shunmin Zhu, Xinbing Wang
IMC6
2024 Canal Mesh: A Cloud-Scale Sidecar-Free Multi-Tenant Service Mesh Architecture
abstract
In recent years, service mesh frameworks have gained significant popularity in building microservice-based applications. A key component of these frameworks is a proxy in each K8s pod, named sidecar, which handles inter-pod traffic. Our empirical measurement reveals that such per-pod sidecars cause numerous problems, including intrusion into the user pod, excessive resource occupation, significant overhead in managing many sidecars, and performance degradation caused by passing traffic through the sidecar.
Enge Song, Yang Song 0031, Chengyun Lu, Tian Pan 0001, Shaokai Zhang, Jianyuan Lu, Jiangu Zhao, Xining Wang, Minglan Gao, Zongquan Li, Ziyang Fang, Biao Lyu, Rong Wen, Li Yi 0003, Zhigang Zong, Shunmin Zhu
SIGCOMM2
2024 CyberStar: Simple, Elastic and Cost-Effective Network Functions Management in Cloud Network at Scale
Bengbeng Xue, Yang Song 0031, Xiaoxin Peng, Yilong Lyu, Xiaoliang Wang 0001, Chen Tian 0001, Cam-Tu Nguyen, Biao Lyu, Rong Wen, Zhigang Zong, Shunmin Zhu
USENIX ATC3
2014 TFA: A Tunable Finite Automaton for Pattern Matching in Network Intrusion Detection Systems
abstract
Deterministic finite automatons (DFAs) and nondeterministic finite automatons (NFAs) are two typical automatons used in the network intrusion detection system. Although they both perform regular expression matching, they have quite different performance and memory usage properties. DFAs provide fast and deterministic matching performance but suffer from the well-known state explosion problem. NFAs are compact, but their matching performance is unpredictable and with no worst case guarantee. In this paper, we propose a new automaton representation of regular expressions, called tunable finite automaton (TFA), to deal with the DFAs' state explosion problem and the NFAs' unpredictable performance problem. Different from a DFA, which has only one active state, a TFA allows multiple concurrent active states. Thus, the total number of states required by the TFA to track the matching status is much smaller than that required by the DFA. Different from an NFA, a TFA guarantees that the number of concurrent active states is bounded by a bound factor b that can be tuned during the construction of the TFA according to the needs of the application for speed and storage. Simulation results based on regular expression rule sets from Snort and Bro show that, with only two concurrent active states, a TFA can achieve significant reductions in the number of states and memory usage, e.g., a 98% reduction in the number of states and a 95% reduction in memory space.
Yang Xu 0010, Junchen Jiang, Rihua Wei, Yang Song 0031, H. Jonathan Chao
IEEE J. Sel. Areas Commun.4