Enge Song

dblp:243/2996 · DBLP profile ↗
← Back
35ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0002-4442-2817ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 29 · 5 first-author · 26 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Scaling LLM Agent Tool Access at Cloud Scale
abstract
LLM agents increasingly rely on tool calling, and the Model Context Protocol (MCP) standardizes it between agents and tool providers, reducing integration cost and driving rapid growth in tool scale. Yet a standardized interface does not make tool access work at production scale: legacy services are not MCP-callable, fast protocol evolution creates compatibility cost, large tool sets exhaust the context window, and stateful sessions complicate load balancing. We solve these with a shared control point, a centralized MCP Gateway System that makes MCP operational at cloud scale. The gateway breaks the direct-connect data plane and consolidates legacy API integration, protocol bridging, access control, and session-aware routing, while scaling out elastically at low per-call overhead. It scales agent tool access to thousands of cloud operations.
Enge Song, Yueshang Zuo, Rong Wen, Jing Tie, Zhou Shao, Qiang Fu 0011, Xiaobo Xue, Luyao Zhong, Shaokai Zhang, Jiangu Zhao, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Changgang Zheng, Tian Pan 0001, Yang Song 0031, Xing Li 0007, Biao Lyu, Meng Li 0010, Haipeng Dai 0001, Guihai Chen, Shunmin Zhu
APNet2
2026 Integrating AI Clusters into Virtual Private Cloud
abstract
While commodity NIC-based back-end AI networks offer ultra-high intra-cluster bandwidth for distributed training, their limited programmability and on-chip resources hinder the implementation of advanced VPC features such as fine-grained isolation and stateful security policies. Furthermore, access to resources within the VPC needs to be routed through the front-end DPU, which is shared by the scale-up domain. The mismatch between the front-end DPU’s bandwidth and the back-end requirements causes GPU underutilization when intensive VPC communication is required for content recommendation, AIGC, and federated learning workloads. We propose an architecture that decouples complex policy enforcement from high-speed packet forwarding to support VPC semantics on back-end NICs and enable front-end/back-end integration. Evaluations show near-full GPU utilization in our analytical model and 71 μ s P999 extra latency of the first packet, suggesting that commodity hardware can support both high-throughput AI training and flexible VPC features.
Xing Li 0007, Enge Song, Changgang Zheng, Shengyao Gao, Juncheng Xiang, Junnan Cai, Haoxiang Pan, Yang Song 0031, Yilong Lv, Qiang Fu 0011, Zhigang Zong, Shunmin Zhu
APNet3
2026 Zephyr: A Zero-loss and Tranparent TLS Connection Migration Framework
abstract
While essential for stateful modern workloads like Large Language Model agents and IoT services, long-lived connections impede cloud infrastructure agility by complicating maintenance and load balancing. Existing connection migration solutions either lack support for industrial-grade encrypted traffic or fail to prevent packet loss during handover in active production environments. To address this gap, we propose Zephyr, a zero-loss and transparent TLS connection migration framework for cross-node migration between servers with different addresses. Zephyr ensures transport-layer consistency by orchestrating an eBPF-based packet buffering mechanism to safely intercept in-flight data. At the application layer, rather than deeply modifying standard TLS libraries, Zephyr creatively reuses the native session resumption mechanism via a “fake client” strategy to reconstruct complex cryptographic states without client involvement. Implemented in widely-used industrial stacks (Nginx and OpenSSL), Zephyr achieves connection migration with approximately 4.1 ms downtime and strict zero packet loss. This approach enables seamless infrastructure optimization without disrupting cloud services.
Chengcheng Yu, Yueshang Zuo, Enge Song, Shaokai Zhang, Jiangu Zhao, Tian Pan 0001, Yang Song 0031, Xing Li 0007, Rong Wen, Chengkun Wei, Shunmin Zhu, Wenzhi Chen
APNet4
2026 Single-Core Hotspots on Your VNF? Break Them Up!
abstract
Current NFVs assign packets to CPU cores at flow granularity, where each flow is pinned to a single CPU. This approach is efficient under most scenarios but has exposed limitations when handling elephant flows. These “heavy hitters” overwhelm single cores, creating bottlenecks that affect overall throughput and degrade service quality. As networks scale to higher-speed links and core-rich CPUs, these imbalances become more severe. In this paper, we propose ParaFlowO, an architecture that Parallelizes processing elephant Flows across multiple CPU cores while preserving in-Order delivery. ParaFlowO breaks elephant flows into flowlets and dynamically rotates them across multiple cores. It integrates a lightweight reordering mechanism to preserve packet order and controls parallelism to mitigate contention on shared state. Preliminary evaluations show that ParaFlowO offers a practical solution to mixed-grained parallelism in stateful middleboxes.
Changgang Zheng, Jin Ke 0005, Enge Song, Yilong Lv, Yisong Qiao, Donglin Lai, Bengbeng Xue, Yang Song 0031, Xing Li 0007, Rong Wen, Zhigang Zong, Shunmin Zhu
APNet5
2026 Bifrost: Alibaba's Next-Generation VPC Network with High-Performance Multipath Reliable Transport
Xing Li 0007, Bo Jiang 0003, Yilong Lv, Yuke Hong, Yinian Zhou, Junnan Cai, Jiayue Xu, Yunrui Hu, Zhao Gao, Enge Song, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Changgang Zheng, Yang Song 0031, Biao Lyu, Rong Wen, Zhigang Zong, Shunmin Zhu
NSDI16
2026 CStar Gateway: Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Tian Pan 0001, Jin Ke 0005, Baohai Hu, Changgang Zheng, Enge Song, Donglin Lai, Yisong Qiao, Bengbeng Xue, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Yang Song 0031, Xionglie Wei, Biao Lyu, Rong Wen, Zhigang Zong, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu
NSDI6
2026 ZooRoute: Enhancing Cloud-Scale Network Reliability via Candidate Path Provisioning and Overlay Proactive Rerouting
Xiaoqing Sun, Xing Li 0007, Xionglie Wei, Tian Pan 0001, Yi Wang 0004, Chenhao Jia, Zhanlong Zhang, Xiaobo Xue, Jianyuan Lu, Shize Zhang, Enge Song, Yang Song 0031, Rong Wen, Biao Lyu, Yang Xu 0010, Shunmin Zhu
NSDI17
2025 Understanding the Long Tail Latency of TCP in Large-Scale Cloud Networks
Enge Song, Bo Jiang 0003, Yang Song 0031, Yuke Hong, Yilong Lv, Yinian Zhou, Junnan Cai, Chao Wang 0128, Yi Wang 0004, Yehao Feng, Shize Zhang, Xiaoqing Sun, Jianyuan Lu, Xing Li 0007, Biao Lyu, Zhigang Zong, Shunmin Zhu
APNet2
2025 Augmenting Public Cloud Infrastructure for Heterogeneous Network Function Virtualization
Yang Song 0031, Tian Pan 0001, Zhigang Zong, Bengbeng Xue, Xionglie Wei, Yisong Qiao, Donglin Lai, Baohai Hu, Jin Ke 0005, Enge Song, Jianyuan Lu, Xing Li 0007, Biao Lyu, Rong Wen, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu
APNet11
2025 Hermes: Enhancing Layer-7 Cloud Load Balancers with Userspace-Directed I/O Event Notification
abstract
Layer-7 load balancers (L7 LBs) improve service performance, availability, and scalability in public clouds. They rely on I/O event notification mechanisms such as epoll to dispatch connections from the kernel to userspace workers. However, early epoll versions suffered from the thundering herd problem. Epoll exclusive (available since Linux 4.5) mitigates this but introduces LIFO wakeups, causing connection concentration on a few workers. Reuseport (Linux 3.9) hashes connections across workers but suffers from hash collisions and lacks awareness of worker load. Since each worker serves multi-tenant traffic, inter-worker load balancing is critical to avoid worker overload and preserve tenant performance isolation.
Tian Pan 0001, Enge Song, Yueshang Zuo, Shaokai Zhang, Yang Song 0031, Jiangu Zhao, Wengang Hou, Jianyuan Lu, Xiaoqing Sun, Shize Zhang, Jiao Zhang 0002, Tao Huang 0005, Biao Lyu, Xing Li 0007, Rong Wen, Zhigang Zong, Shunmin Zhu
SIGCOMM2
2025 Nezha: SmartNIC-based Virtual Switch Load Sharing
abstract
Cloud providers use SmartNIC-accelerated virtual switches (vSwitches) to offer rich network functions (NFs) for tenant VMs. Constrained by limited SmartNIC resources, it is a challenge to provide sufficient network performance for high-demand VMs. Meanwhile, we observed a significant number of idle vSwitches in the data center, which led us to consider leveraging them to build a remote resource pool for high-demand virtual NICs (vNICs). In this work, we propose Nezha, a distributed vSwitch load sharing system. Nezha reuses the existing idle SmartNICs to handle the excess load from the local SmartNIC without adding new devices. Nezha offloads stateless rule/flow tables to the remote, while keeping states locally. This eliminates the need for state synchronization, facilitating load sharing and failover. The deployment cost of Nezha is only a small fraction of that required to deploy new devices. Data collected from production show that our CPS capability bottleneck has shifted from the vSwitch to the VM kernel stack, with #concurrent flows and #vNICs increased by up to 50.4x and 40x, respectively.
Xing Li 0007, Enge Song, Tian Pan 0001, Qiang Fu 0011, Yang Song 0031, Yilong Lv, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Rong Wen, Xionglie Wei, Biao Lyu, Zhigang Zong, Qinming He, Shunmin Zhu
SIGCOMM2
2025 Albatross: A Containerized Cloud Gateway Platform with FPGA-accelerated Packet-level Load Balancing
abstract
Alibaba Cloud's centralized gateways relied heavily on high-capacity switching ASICs, but the abrupt halt of Tofino chip evolution in Jan 2023 forced us to seek alternatives that can meet the requirements of performance, supply-chain security, code reuse, and resource efficiency. After evaluating multiple options, we developed Albatross, our 3rd gen cloud gateway based on FPGA and x86 CPUs. Albatross delivers FPGA-based packet-level load balancing to the host CPUs to prevent CPU core overload, manages large reorder buffers under high-latency jitters (100μs) during complex cloud service processing, and resolves head-of-line (HOL) blocking from packet losses or software exceptions in CPUs. To avoid being overloaded by heavy hitters due to anomalies or attacks, it also implements a two-stage rate limiter for millions of tenants with only 2MB of FPGA memory. To maximize resource utilization, Albatross uses containerization to host multiple gateway instances and designs a BGP proxy to lessen the BGP peering overhead on uplink switches caused by high-density container deployments. After hundreds of man-months of development, a single Albatross node can process 80~120Mpps of cloud network traffic with an average latency of 20μs, reducing gateway and sandbox infra costs by 50%.
Jianyuan Lu, Shunmin Zhu, Tian Pan 0001, Yisong Qiao, Yang Song 0031, Wenqiang Su, Yanqiang Li, Enge Song, Shize Zhang, Xiaoqing Sun, Rong Wen, Xionglie Wei, Biao Lyu, Xing Li 0007
SIGCOMM11
2025 ZooRoute: Enhancing Cloud-Scale Network Reliability via Overlay Proactive Rerouting
abstract
This paper presents ZooRoute, a tenant-transparent, fast failure recovery service that requires no modifications to physical devices. ZooRoute leverages the overlay layer and enables traffic flows to bypass failures by altering source ports (srcPorts) in packet headers during encapsulation. To enable deployment in large-scale cloud networks, ZooRoute proposes: 1) On-demand probing to efficiently monitor a vast number of hosts while minimizing telemetry costs. 2) Table compression to record the states of numerous paths with limited on-chip resources. 3) A device-sensing mechanism to prevent unnecessary reconnections in stateful forwarding. Deployed in Alibaba Cloud for 18 months, ZooRoute has significantly improved network reliability, reducing cumulative outage time by 92.71%.
Xiaoqing Sun, Xionglie Wei, Xing Li 0007, Yi Wang 0004, Chenhao Jia, Zhanlong Zhang, Jianyuan Lu, Shize Zhang, Enge Song, Yang Song 0031, Tian Pan 0001, Rong Wen, Biao Lyu, Yang Xu 0010, Shunmin Zhu
SIGCOMM15
2024 QuarkTable: Building Compact Forwarding Tables for Programmable Switches on Public Clouds
abstract
Programmable switches have been recently proposed as dataplane solutions for public clouds. However, the conflict of limited on-chip memory and massive forwarding rules in cloud networks hinders the large-scale deployments. We argue that building compact forwarding tables for programmable switches is a viable option to this problem. In this paper, as a first step, we explore the feasibility of compact data structures for VPC routing tables (VRTs) and propose QuarkTable as a solution. The idea of QuarkTable builds upon the existence of redundancy in the prefixes of VRTs, supported by extensive analysis of real-world VRTs collected from six geographically distributed regions of Alibaba Cloud. By cutting the VRT into two partitions and encoding the upper part of the longer prefixes, QuarkTable effectively shortens the length of each entry and reduces the overall memory consumption. We demonstrate the effectiveness of QuarkTable by showing its proximity to the entropy bound of real-world VRTs. Experiments on six VRTs from Alibaba Cloud show practical memory savings up to 33.5% and 30.2% for SRAM and TCAM, respectively.
Jianyuan Lu, Huaiyi Zhao, Yehao Feng, Shengru Li, Enge Song, Xionglie Wei, Biao Lyu, Rong Wen, Shunmin Zhu
APNet7
2024 vSwitchLB: Stratified Load Balancing for vSwitch Efficiency in Data Centers
abstract
The virtual switch (vSwitch) serves as a fundamental element in cloud network, critical for high-performance and strongly isolated inter-VM forwarding in local and external networks. Similar to other multicore systems, a vSwitch with multiple cores also faces the issue of core load imbalance. As a major cloud provider, we pinpoint four cases of core load imbalance within the vSwitch in our cloud, stemming from unequal traffic distribution across virtual queues and RSS buckets, as well as from traffic patterns like heavy hitters and micro-bursts. To tackle the different load imbalance cases, we present vSwitchLB, a vSwitch load balance framework. Specifically, we introduce a load imbalance detection module, accompanied by dedicated techniques designed to address each specific type of imbalance. Our preliminary evaluation shows that vSwitchLB can accurately classify different load imbalances encountered in the vSwitch on our cloud and then prevent any single core of vSwitch from being flooded and overwhelmed.
Enge Song, Yi Wang 0004, Jianyuan Lu, Xing Li 0007, Biao Lyu, Rong Wen, Shibo He, Yuanchao Shu, Shunmin Zhu
APNet2
2024 CloudPlanner: Minimizing Upgrade Risk of Virtual Network Devices for Large-Scale Cloud Networks
abstract
Cloud networks continuously upgrade softwarized virtual network devices (VNDs) to meet evolving tenant demands. However, such upgrades may result in unexpected failures. An intuitive idea to prevent upgrade failures is to resolve all compatibility issues before deployment, but it is impractical to replicate all deployed VND cases and test them with lots of replayed real traffic for the VND developers. As a result, the operations team takes upgrade risk to test upgrades by gradually deploying them. Although careful upgrade schedule planning is the most common method to minimize upgrade risk, to the best of our knowledge, no VND upgrade schedule planning scheme has been adequately studied for large-scale cloud networks. To fill this gap, we propose CloudPlanner, the first VND upgrade schedule planning scheme aiming to minimize the VND upgrade risk for large-scale cloud networks. CloudPlanner prioritizes upgrading VNDs that are more likely to trigger failures based on expert knowledge and historical failure-trigger VND properties and limits the number of tenants associated with simultaneously upgraded VNDs. We also propose a heuristic solver which can quickly and greedily plan schedules. Using real-world data from production environments, we demonstrate the benefits of CloudPlanner through extensive experiments.
Enhuan Dong, Jiahai Yang 0001, Shize Zhang, Zejie Wang, Xiaoqing Sun, Enge Song, Jianyuan Lu, Biao Lyu, Shunmin Zhu
INFOCOM10
2024 LuoShen: A Hyper-Converged Programmable Gateway for Multi-Tenant Multi-Service Edge Clouds
Tian Pan 0001, Xionglie Wei, Yisong Qiao, Tiesheng Cheng, Wenqiang Su, Yuke Hong, Zhengzhong Wang, Chongjing Dai, Peiqiao Wang, Xuetao Jia, Jianyuan Lu, Enge Song, Biao Lyu, Ennan Zhai, Jiao Zhang 0002, Tao Huang 0005, Dennis Cai, Shunmin Zhu
NSDI18
2024 POSEIDON: A Consolidated Virtual Network Controller that Manages Millions of Tenants via Config Tree
Biao Lyu, Enge Song, Tian Pan 0001, Jianyuan Lu, Shize Zhang, Xiaoqing Sun, Chenxiao Wang, Xiuheng Chen, Yandong Duan, Weisheng Wang, Jinpeng Long, Kunpeng Zhou, Zhigang Zong, Xing Li 0007, Guangwang Li, Peng Cheng 0001, Jiming Chen 0001, Shunmin Zhu
NSDI2
2024 Canal Mesh: A Cloud-Scale Sidecar-Free Multi-Tenant Service Mesh Architecture
abstract
In recent years, service mesh frameworks have gained significant popularity in building microservice-based applications. A key component of these frameworks is a proxy in each K8s pod, named sidecar, which handles inter-pod traffic. Our empirical measurement reveals that such per-pod sidecars cause numerous problems, including intrusion into the user pod, excessive resource occupation, significant overhead in managing many sidecars, and performance degradation caused by passing traffic through the sidecar.
Enge Song, Yang Song 0031, Chengyun Lu, Tian Pan 0001, Shaokai Zhang, Jianyuan Lu, Jiangu Zhao, Xining Wang, Minglan Gao, Zongquan Li, Ziyang Fang, Biao Lyu, Rong Wen, Li Yi 0003, Zhigang Zong, Shunmin Zhu
SIGCOMM1
2024 CloudSentry: Two-Stage Heavy Hitter Detection for Cloud-Scale Gateway Overload Protection
abstract
The cloud vendors provide sharing resources for millions of tenants across the world to achieve economies of scale. At the same time, the cloud network keeps the performance isolation between different tenants as if they use their private dedicated resources. However, heavy hitters caused by a single tenant at cloud gateways will break such isolation, undermining the predictable performance expected by other cloud tenants. To prevent it, heavy hitter detection becomes a key concern at the performance-critical cloud gateways but faces the dilemma between fine granularity and low overhead. In this work, we presentCloudSentry, a scalable two-stage heavy hitter detection system dedicated to multi-tenant cloud gateways against such a dilemma. CloudSentry uses CPU utilization as an indicator of heavy hitters and conducts a lightweight coarse-grained detection running 24/7 to detect such CPU spikes. Then it invokes a fine-grained detection to precisely dump and analyze the potential heavy-hitter packets at the CPU spikes. After that, a more comprehensive analysis is conducted to associate heavy hitters with the cloud service scenarios and invoke a corresponding backpressure procedure. CloudSentry significantly reduces memory, computation and storage overhead compared with existing approaches. In a gateway cluster under an average traffic throughput of 251 Gbps, CloudSentry consumes only a fraction of 2%–5% CPU utilization with 8 KB run-time memory, producing only 10 MB heavy hitter logs during one month. Additionally, as it has been deployed in Alibaba Cloud for over two years, we share case studies and a lot of deployment experiences in this article.
Jianyuan Lu, Tian Pan 0001, Mao Miao, Guangzhe Zhou, Yining Qi, Shize Zhang, Enge Song, Xiaoqing Sun, Huaiyi Zhao, Biao Lyu, Shunmin Zhu
IEEE Trans. Parallel Distributed Syst.8
2024 INT-Label: Lightweight In-Band Network-Wide Telemetry via Distributed Labeling
abstract
In-band Network Telemetry (INT) enables hop-by-hop device-internal state exposure for maintaining and troubleshooting data center networks. To achievenetwork-widetelemetry coverage, orchestration on top of the INT primitive is required. A straightforward solution would flood the network with INT probe packets for maximum measurement coverage, which leads to a huge bandwidth overhead. A refined solution leverages the SDN controller to collect the network topology information and carry out centralized probing path planning, which, however, is inefficient in reacting to topology changes. To tackle the above problems, we proposeINT-label, a lightweight In-band Network-Wide Telemetry architecture via the distributed labeling approach. INT-label periodically labels the sampled packets with device-internal states. It is cost-effective with a minor bandwidth overhead and able to seamlessly adapt to topology changes. In order to reduce the number of labeled packets, we introduce a times-based probabilistic labeling algorithm, which allows fewer packets to carry more INT information than the interval-based algorithm. In addition, to counteract the degradation of telemetry resolution due to loss of labeled packets, we design a feedback mechanism which can adaptively change the instant labeling frequency. We provide theoretical proof that INT-label can achieve network-wide telemetry. We analyze the impact of transmission delay on coverage rate and labeling times distribution under the INT-label architecture. Evaluation on software P4 switches suggests that INT-label can achieve 99.72% measurement coverage under the labeling frequency of 20 times per second. With the adaptive labeling enabled, even if 60% of the packets are lost, the coverage can still reach 92%.
Enge Song, Tian Pan 0001, Haoyu Song 0001, Qiang Fu 0011, Yingjiang Liu, Chenhao Jia, Chuanying Yuan, Minglan Gao, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001
IEEE Trans. Parallel Distributed Syst.1
2023 INT-Balance: In-Band Network-Wide Telemetry with Balanced Monitoring Path Planning
abstract
In-band Network Telemetry (INT) empowers high-resolution network monitoring by collecting hop-by-hop device-internal states through the data plane without frequently disturbing the control plane. To achieve network-wide monitoring, a high-level orchestration is made to provision multiple monitoring paths to cover the entire network. The path number and path overlapping are kept minimum to maximally reduce the telemetry overhead. However, in production deployment, except for the telemetry overhead, the telemetry timeliness is equally important for fine-grained monitoring, which creates new requirements of balanced monitoring path planning. Given the INT probes from multiple paths are collected to the central controller for analysis, the late arrival of even one probe will delay the analysis process and affect the monitoring timeliness. To address the problem, we propose INT-balance, a novel path planning algorithm for balanced INT path generation. In INT-balance, we first break the original network graph into multiple path segments at the odd vertices. Then, we iteratively splice the two shortest path segments with the joint endpoints to form a longer path segment until the path segment number reaches half the number of the odd vertices. INT-balance generates the minimum number of INT paths with well-balanced path lengths, covering every edge of the network graph without any path overlapping. Evaluation on a network of 100 switches shows that the path length variance of INT-balance is 67% less than that of INT-path, while the algorithm execution time is increased only by 0.012s.
Yan Zhang 0063, Tian Pan 0001, Enge Song, Jiang Liu 0010, Tao Huang 0005, Yunjie Liu 0001
ICC4
2022 Power-Aware Traffic Engineering for Data Center Networks via Deep Reinforcement Learning
abstract
The issue of high energy consumption and low energy utilization in data center networks (DCNs) has always been the focus of attention of both academia and industry. One general solution is to select a subset of network devices that can meet the traffic transmission requirements, thereby turning off the remaining redundant devices. However, modeling the problem as integer linear programming introduces significant time overhead, while heuristic approaches often suffer from poor generalizability. In this paper, we propose GreenDCN.ai, a closed-loop control system, which utilizes In-band Network Telemetry to collect the network-wide device-internal state, and leverages a Deep Reinforcement Learning-based energy-saving algorithm to make rapid decisions to turn on or off network device ports in response to the real-time network state. The trained GreenDCN.ai can adaptively adjust its energy-saving strategy without human intervention when the DCN topology changes. Besides, based on the regularity of the DCN topology, we design two training complexity reduction methods to address the non-convergence issue under large-scale DCN topologies. Specifically, we split the large-scale DCN topology into sub-topologies for parallel training on each sub-topology without breaking the DCN topology connectivity. Evaluation on software P4 switches suggests that GreenDCN.ai can achieve stable convergence within 590 episodes, generate effective action decisions within$\boldsymbol{79}\upmu\mathrm{s}$, and save about 34% to 39% of the network energy consumption.
Minglan Gao, Tian Pan 0001, Enge Song, Mengqi Yang, Tao Huang 0005, Yunjie Liu 0001
GLOBECOM3
2022 MIMIC: SmartNIC-aided Flow Backpressure for CPU Overloading Protection in Multi-Tenant Clouds
abstract
In multi-tenant clouds, off-the-shelf x86 boxes are widely deployed as middleboxes. With the rapid growth of cloud traffic and the migration to NFV deployment in recent years, CPU overloading at middleboxes becomes more of an issue. From our data centers, we observed that the CPU overloading was caused by heavy hitters. To address this issue, we propose MIMIC, a cloud-scale flow backpressure system, implemented onto our existing SmartNIC with FPGA acceleration. MIMIC rate-limits the selected heavy hitters through a new per-flow backpressure protocol and a new heavy-hitter detection system, to protect the other tenants. The detection system is based on hierarchical memory design, leveraging on-chip SRAM and off-chip DRAM, which can handle highly concurrent cloud traffic without the losses of flow information. We extend the design by adding a pre-filtering procedure for rapid detection. To avoid CPU being flooded by FPGA through frequent heavy-hitter reporting, due to their performance disparity, the CPU queries the FPGA on demand. The backpressure protocol is non-invasive to protect tenant privacy and allows controllable rate-limiting through the novel use of ECN and meter tables. The SmartNIC acts as a man in the middle to facilitate heavy-hitter detection and per-flow backpressuring. In a production setting, we observe that MIMIC can react quickly and bring down CPU load to the normal level within 10ms without packet losses.
Enge Song, Nianbing Yu, Tian Pan 0001, Qiang Fu 0011, Xionglie Wei, Yisong Qiao, Jianyuan Lu, Yijian Dong, Mingxu Xie, Jinkui Mao, Zhengjie Luo, Chenhao Jia, Jiao Zhang 0002, Tao Huang 0005, Biao Lyu, Shunmin Zhu
ICNP1
2022 WebQMon.ai: Gateway-Based Web QoE Assessment Using Lightweight Neural Networks
Enge Song, Tian Pan 0001, Qiang Fu 0011, Chenhao Jia, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001
ICSOC1
2021 A Refined Dijkstra's Algorithm with Stable Route Generation for Topology-Varying Satellite Networks
abstract
SpaceX plans ambitiously to launch approximately 12,000 satellites from 2019 to 2024, expected to be a complement or even competitor to ground networks. However, the mega-scale satellite network is topology-varying and the frequency of inter-satellite link (ISL) handovers increases rapidly as the topology expands, which will further arouse a massive number of route updates with considerable packet travel delay or even packet loss during the route convergence. The classic Dijkstra's algorithm is adopted for space route calculation, however, it always selects the default shortest path from multiple equal-cost shortest paths between two satellite nodes. To reduce the route change as much as possible during the periodical topology change, in this work, we refined the original Dijkstra and propose StableRoute to select the most appropriate route from the equal-cost candidates with the least route updates compared with the routing table last round. In this way, the end-to-end paths can be maintained as far as possible without time-to-time oscillation. Evaluation shows that it reduces 41% of the route updates in a 36 × 36 topology compared with Dijkstra, and the reduction rate will rise persistently with the growth of the satellite constellation.
Zhengjie Luo, Tian Pan 0001, Enge Song, Houtian Wang, Wenhao Xue, Tao Huang 0005, Yunjie Liu 0001
ICDCS3
2021 INT-probe: Lightweight In-band Network-Wide Telemetry with Stationary Probes
abstract
Visibility is essential for operating and troubleshooting intricate networks. In-band Network Telemetry (INT) has been embedded in the latest merchant silicons to offer high-precision device and traffic state visibility. INT is actually an underlying technique and each INT instance covers only one monitoring path. The network-wide measurement coverage therefore requires a high-level orchestration to provision multiple INT paths. An optimal path planning is expected to produce a minimum number of paths with a minimum number of overlapping links. Eulerian trail has been used to solve the general problem. However, in production networks, the vantage points where one can deploy probes to start and terminate INT paths are constrained. In this work, we propose an optimal path planning algorithm, INT-probe, which achieves the network-wide telemetry coverage under the constraint of stationary probes. INT-probe formulates the constrained path planning into an extended multi-depot k-Chinese postman problem (MDCPP-set) and then reduces it to a solvable minimum weight perfect matching problem. We analyze algorithm's theoretical bound and the complexity. Extensive evaluation on both wide area networks and data center networks with different scales and topologies are conducted. We show INT-probe is efficient, high-performance, and practical for real-world deployment. For a large-scale data center networks with 1125 switches, INT-probe can generate 112 monitoring paths (reduced by 50.4 %) by allowing only 1.79% increase of the total path length, promptly resolving link failures within 744.71ms.
Tian Pan 0001, Xingchen Lin, Haoyu Song 0001, Enge Song, Zizheng Bian, Hao Li 0011, Jiao Zhang 0002, Fuliang Li, Tao Huang 0005, Chenhao Jia, Bin Liu 0001
ICDCS4
2021 INT-label: Lightweight In-band Network-Wide Telemetry via Interval-based Distributed Labelling
abstract
The In-band Network Telemetry (INT) enables hop-by-hop device-internal state exposure for reliably maintaining and troubleshooting data center networks. For achieving network-wide telemetry, orchestration on top of the INT primitive is further required. One straightforward solution is to flood the INT probe packets into the network topology for maximum measurement coverage, which, however, leads to huge bandwidth overhead. A refined solution is to leverage the SDN controller to collect the topology and carry out centralized probing path planning, which, however, cannot seamlessly adapt to occasional topology changes. To tackle the above problems, in this work, we propose INT-label, a lightweight In-band Network-Wide Telemetry architecture via interval-based distributed labelling. INT-label periodically labels device-internal states onto sampled packets, which is cost-effective with minor bandwidth overhead and able to seamlessly adapt to topology changes. Furthermore, to avoid telemetry resolution degradation due to loss of labelled packets, we also design a feedback mechanism to adaptively change the instant label frequency. Evaluation on software P4 switches suggests that INT-label can achieve 99.72% measurement coverage under a label frequency of 20 times per second. With adaptive labelling enabled, the coverage can still reach 92% even if 60% of the packets are lost in the data plane.
Enge Song, Tian Pan 0001, Chenhao Jia, Wendi Cao, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001
INFOCOM1
2021 GreenTE.ai: Power-Aware Traffic Engineering via Deep Reinforcement Learning
abstract
Power-aware traffic engineering via coordinated sleeping is usually formulated into Integer Programming problems, which are generally NP-hard with unbounded computation time for large-scale networks. This results in delayed control decision making in dynamic network environments. Motivated by advances in deep Reinforcement Learning, we consider building intelligent systems that learn to adaptively change router/switch’s power state according to changing network conditions. Neural network’s forward propagation can greatly speed up power on/off decision making. Generally, conducting RL requires a learning agent to iteratively explore and perform the "good" actions based on the feedback from the environment. By coupling Software-Defined Networking for performing centrally calculated actions to the environment and In-band Network Telemetry for collecting feedback from the environment, we develop GreenTE.ai, a closed-loop control/training system to automate power-aware traffic engineering. Furthermore, we propose novel techniques to enhance the learning ability and reduce the learning complexity. With both energy efficiency and traffic load balancing considered, GreenTE.ai can generate reasonable power saving actions within 276ms under a network testbed of 11 software P4 switches.
Tian Pan 0001, Xiaoyu Peng, Zizheng Bian, Xingchen Lin, Enge Song, Fuliang Li, Yang Xu 0010, Tao Huang 0005
IWQoS6
2021 Sailfish: accelerating cloud-scale multi-tenant multi-service gateways with programmable switches
abstract
The cloud gateway is essential in the public cloud as the central hub of cloud traffic. We show that horizontal scaling of software gateways, once sustainable for years, is no longer future-proof facing the massive scale and rapid growth of today's cloud. The root cause is the stagnant performance of the CPU core, which is prone to be overloaded by heavy hitters as traffic growth goes far beyond Moore's law. To address this, we propose \emph{Sailfish}, a cloud-scale multi-tenant multi-service gateway accelerated by programmable switches. The new challenge is that large forwarding tables due to multi-tenancy cannot be fit into the limited on-chip memories. To this end, we devise a multi-pronged approach with (1) hardware/software co-design for table sharing, (2) horizontal table splitting among gateway clusters, (3) pipeline-aware table compression for a single node. Compared with the x86 gateway of a similar price, Sailfish reduces latency by 95% (2μs), improves throughput by more than 20x in bps (3.2Tbps) and 71x in pps (1.8Gpps) with packet length < 256B. Sailfish has been deployed in Alibaba Cloud for more than two years. It is the first P4-based cloud gateway in the industry, of which a single cluster carries dozens of Tbps traffic, withstanding peak-hour traffic in large online shopping festivals.
Tian Pan 0001, Nianbing Yu, Chenhao Jia, Jianwen Pi, Yisong Qiao, Jianyuan Lu, Enge Song, Jiao Zhang 0002, Tao Huang 0005, Shunmin Zhu
SIGCOMM11
2021 NB-Cache: Non-Blocking In-Network Caching for High-Performance Content Routers
abstract
Information-Centric Networking (ICN) provides scalable and efficient content distribution at the Internet scale due to in-network caching and native multicast. To support these features, a content router needs high performance at its data plane, which consists of three forwarding steps: checking the Content Store (CS), then the Pending Interest Table (PIT), and finally the Forwarding Information Base (FIB). In this work, we build an analytical model of the router and identify that CS is the actual bottleneck. Then, we propose a novel mechanism called “NB-Cache” to address CS’s performance issue from a network-wide point of view. In NB-Cache, when packets arrive at a router whose CS is fully loaded, instead of being blocked and waiting for the CS, these packets are forwarded to the next-hop router, whose CS may not be fully loaded. This approach essentially utilizes Content Stores of all the routers along the forwarding path in parallel rather than checking each CS sequentially. NB-Cache follows a design pattern of on-demand load balancing and can be formulated into a non-trivial N-queue bypass model. We use the Markov chain to establish its theoretical base and find an algorithm for automated transition rate matrix generation. Experiments show significant improvement of data plane performance: 70% reduction in round-trip time (RTT) and 130% increase in throughput. NB-Cache decouples the fast packet forwarding from the slower content retrieval thus substantially reducing CS’s heavy dependency on fast but expensive memory.
Tian Pan 0001, Xingchen Lin, Enge Song, Jiao Zhang 0002, Hao Li 0011, Jianhui Lv, Tao Huang 0005, Bin Liu 0001, Beichuan Zhang 0001
IEEE/ACM Trans. Netw.3
2020 INT-filter: Mitigating Data Collection Overhead for High-Resolution In-band Network Telemetry
abstract
In-band Network Telemetry (INT) enables fine-grained network monitoring to ease the management of large-scale networks, which, however, relies on the real-time collection of a huge amount of telemetry data through the southbound interface. For example, the INT telemetry data upload rate of a 28-pod FatTree topology reaches 3Tbps under a probe frequency of 100 times/s, which is rather unacceptable since the controller-switch link bandwidth is limited. To mitigate the telemetry data collection overhead, in this work, we propose INT-filter, a novel measurement architecture that deploys the same prediction algorithm on both the data plane and the control plane to predict the traffic state in the near future instead of uploading all the telemetry data. Such prediction-based approach leverages the observation that there is considerable redundancy in the telemetry data sequence. In addition, we design an integration mechanism that conducts predictions using multiple methods simultaneously and uploads the predicted result from the least-error method to further decrease the upload volume. Extensive evaluation suggests that INT-filter can achieve at least 33.6% data collection decrease under a 10ms probe interval. With prediction integration, the upload reduction can further reach 58.5%.
Enge Song, Tian Pan 0001, Chenhao Jia, Wendi Cao, Jiao Zhang 0002, Tao Huang 0005, Yunjie Liu 0001
GLOBECOM1
2020 Rapid Detection and Localization of Gray Failures in Data Centers via In-band Network Telemetry
abstract
Network reliability becomes increasingly important in modern data center networks (DCNs). The DCNs are expected to work sustainably under internal failures and assist network operators in troubleshooting them rapidly. However, some network failures will happen silently with packets discarded without producing any explicit notification before causing tremendous damage to the network. To troubleshoot these "gray failures", in this work, we present a rapid gray failure detection and localization mechanism based on the recently proposed In-band Network Telemetry (INT). Specifically, we leverage simplified INT probe packets to conduct network-wide telemetry to help the servers under ToR switches obtain all the feasible paths between sources and destinations. Once a network failure occurs, the affected thus unavailable paths will immediately be detected and flushed out of the path information table at each server by a timeout mechanism. Hence, servers can proactively perform source routing-based fast traffic reroute to avoid massive packet loss and retain uninterrupted quality of experience. At the meantime, all the aged path entries will be uploaded to a remote controller for centralized failure localization by identifying common path elements. To verify the feasibility of our design, we build a virtual network testbed with software P4 switches and a Redis database. Evaluation shows that our system can successfully detect network gray failures and reroute the affected traffic in no time while complete failure localization within only a few seconds.
Chenhao Jia, Tian Pan 0001, Zizheng Bian, Xingchen Lin, Enge Song, Tao Huang 0005, Yunjie Liu 0001
NOMS5
2020 Threshold-oblivious on-line web QoE assessment using neural network-based regression model
abstract
The evaluation of the web‐browsing quality of experience (QoE) is difficult to complete through traditional methods (e.g. deducing formulas or setting thresholds) due to the diversity of websites and their contents. To evaluate web‐browsing QoE through a general way, the authors propose a web QoE evaluation architecture based on machine learning, consisting of two parts: traffic classification sub‐system and QoE prediction sub‐system. When evaluating user experience, traffic classification sub‐system first classifies the packets generated by visiting a website into a flowthrough some fields in the packet header, to model each website separately. The traffic classification accuracy of packets over six websites reaches 96.63%. Then, in the network layer, the traffic metric cumulative traffic volume is generated from the size and arrival time of packets. When a user visits a web page, their regression model predicts the above‐the‐fold time (ATF) and thus QoE. The output of the regression model is an exact ATF value that is mapped to user experience. In addition, reversing input variables further improves the model, which is evaluated on two popular websites. The QoE prediction results of the improved method for 5400 visits are obtained within 0.0975 s, reaching 0.9 .
Enge Song, Tian Pan 0001, Qiang Fu 0011, Chenhao Jia, Wendi Cao, Tao Huang 0005
IET Commun.1
2019 INT-path: Towards Optimal Path Planning for In-band Network-Wide Telemetry
abstract
With the ever-increasing complexity of networks, fine-grained network monitoring enables better network reliability and timely feedback control. The In-band Network Telemetry (INT) allows cost-effective network monitoring by encapsulating device-internal states into probe packets. However, INT only specifies an underlying device-level primitive while how to achieve network-wide traffic monitoring remains undefined. In this work, we propose INT-path, a network-wide telemetry framework, by decoupling the system into a routing mechanism and a routing path generation policy. Specifically, we embed source routing into INT probes to allow specifying the route the probe packet takes through the network. Above the mechanism, we develop an Euler trail-based path planning policy to generate non-overlapped INT paths that cover the entire network with a minimum path number. Besides, an exhaustive analysis of algorithm's run-time complexity is also provided. INT-path can “encode” the network-wide traffic status into a series of “bitmap images”, transforming network troubleshooting into pattern recognition problems. INT-path is very suitable for deployment in data center networks thanks to their symmetric network topologies.
Tian Pan 0001, Enge Song, Zizheng Bian, Xingchen Lin, Xiaoyu Peng, Jiao Zhang 0002, Tao Huang 0005, Bin Liu 0001, Yunjie Liu 0001
INFOCOM2