EDBT 2026 Demo / reviewers in the wild / expert
Chen Sun 0005
dblp:01/6072-5
· DBLP profile ↗
39ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0003-2480-2350ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 33 · 7 first-author · 10 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OmniPath Ping: Active Network Measurement in the Era of Packet Spraying
Kaicheng Yang 0001, Zongwei Lv, Peijun Huang, Kaitai Zhang, Qiuheng Yin, Yaoming Li, Feiyu Wang 0002, Zhuochen Fan, Yikai Zhao 0001, Chen Sun 0005, Tong Yang 0003 |
SIGCOMM | 10 |
| 2025 | TraceWizard: End-to-End Distributed Tracing Across Host and Network Devices in CloudabstractThe rise of microservice architecture in cloud computing has introduced additional complexities in diagnosing faults, as traditional end-to-end tracing systems often fail to address issues beyond the application layer, such as network devices. To overcome this limitation, we introduce TraceWizard, an enhanced end-to-end tracing system that integrates eBPF and SDN (Software Defined Network) technologies to track requests across host and network devices. By enabling full life-cycle tracing and maintaining consistent trace contexts, TraceWizard provides fine-grained insights into faults across applications, the OS kernel, and network devices. Our evaluations show that its data enables more effective fault detection, achieving an average accuracy of 91.6 % across different algorithms—significantly outperforming application-layer monitoring tools. Additionally, it helps operators identify root causes with minimal overhead, reducing QPS by 2.2%, increasing QCT by 2.2%, and adding 3.41 % CPU and 2.43% memory usage. Kuangyuan Li, Jingrun Zhang, Pengfei Chen 0002, Hongyang Chen 0002, Ruipeng Hong, Wanqi Yang, Chen Sun 0005 |
CLOUD | 7 |
| 2025 | CounterSnake: A lossless and generalized compression framework for diverse sketches
Xunpeng Liu, Qun Huang 0001, Yaojing Wang, Lihua Miao, Chen Sun 0005 |
Proc. VLDB Endow. | 5 |
| 2025 | Zhuge: Toward Consistent Low Latency With Minimal Control Loop DelayabstractReal-time communication (RTC) applications demand consistent low latency to ensure a smooth and interactive user experience. However, wireless networks, including WiFi and cellular, although they provide satisfactory median latency, often suffer from significant tail latency due to the highly variable network bandwidth. We observe that the control loop for managing the sending rate of RTC applications becomes inflated when congestion occurs at the wireless access point (AP), leading to untimely rate adaptation in response to wireless dynamics. Existing solutions fail to quickly adapt to bandwidth fluctuations due to the inflated control loop. In this paper, we propose Zhuge, a purely wireless AP-based solution that addresses these issues by separating congestion feedback from congested queues. Our approach involves the design of a Fortune Teller, which accurately estimates the wireless latency for each packet upon its arrival at the wireless AP. To ensure scalability, we also develop a Feedback Updater that translates the estimated latency into understandable feedback messages for various end-to-end protocols, delivering them back to the senders immediately for rate adaptation. Our evaluation, based on both trace-driven simulations and real-world scenarios, demonstrates that Zhuge significantly reduces the occurrence of large tail latency and alleviates RTC performance degradation by 22% to 95%. Bo Wang 0066, Xingxing Yang 0008, Zili Meng, Yaning Guo, Chen Sun 0005, Justine Sherry, Hongqiang Harry Liu, Mingwei Xu 0001 |
IEEE Trans. Netw. | 5 |
| 2025 | NetScope: Fault Localization in Programmable Networking Systems With Low-Cost In-Band Network Telemetry and In-Network DetectionabstractRecently, Software Defined Networking (SDN) has gained widespread adoption as a network infrastructure. Although the openness and programmability of SDN facilitate large complex network construction, diagnosing faults in datacenter-scale network remains challenging. Previous network diagnosis tools pose significant overhead in fine-grained telemetry and typically lack automated fine-grained fault diagnosis capabilities. Although on-demand monitoring methods have been proposed to reduce telemetry overhead, they struggle with effectively setting fixed thresholds, which requires expert experience. This paper presents NetScope, a lightweight system for real-time anomaly detection with self-adaptive thresholds and automatic root cause localization in programmable networking systems. NetScope estimates latency medians for each Flow (i.e., a pair of source and sink switches) within the switch using the proposed per-Flow quantile sketch and calculates the threshold accordingly for anomaly detection. Upon detecting anomalies, NetScope collects aggregated packet-level telemetry on demand and generates a ranked list of fine-grained fault culprits at multiple levels, including port-level, Flow-level, and switch-level. Extensive experiments demonstrate the effectiveness and efficiency of NetScope in anomaly detection and fault localization. Specifically, NetScope achieves a 32%~116% relative improvement in anomaly detection and 6%~197% improvement in root cause analysis compared with other baselines without causing any network bandwidth in anomaly detection while consuming 64.2% less telemetry bandwidth for localization. Hongyang Chen 0002, Benran Wang, Guangba Yu, Pengfei Chen 0002, Chen Sun 0005, Zibin Zheng |
IEEE Trans. Netw. | 6 |
| 2023 | Buffer-Based High-Coverage and Low-Overhead Request Event Monitoring in the CloudabstractRequest latency directly affects the performance of modern cloud applications. Due to various causes in hosts and networks, requests can suffer from request latency anomalies (RLAs), which may violate the Service-Level Agreement. However, existing performance monitoring tools have incomplete coverage and inconsistent semantics for monitoring requests and cannot accurately diagnose RLAs. This paper presentsBufScope, a high-coverage and low-overhead request event monitoring system, which monitorsbuffersto capture most RLA-related abnormal events with consistent request-level semantics in the end-to-end datapath of request. First,BufScopemodels the datapath of request as a buffer chain and defines events based on three properties of buffers, so as toend-to-end monitorthe root causes of RLA. Then, to achieveconsistent semanticsfor captured events,BufScopedesigns a request-level semantics injection mechanism to make events captured in networks have the victim requests’ ID. Finally,BufScopeoffloads the semantics operations and event collection in software to SmartNICs forlow CPU overhead. We have implementedBufScopeon commodity SmartNICs and programmable switches. Evaluation results show thatBufScopecan diagnose 98% RLAs with < 0.08% network bandwidth overhead and 0.6% application throughput decline. Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005, Lu Lu 0016 |
IEEE/ACM Trans. Netw. | 2 |
| 2023 | Dependable Virtualized Fabric on Programmable Data PlaneabstractIn modern multi-tenant data centers, each tenant desires reassuring dependability from the virtualized network fabric – bandwidth guarantee with work conservation, bounded tail latency and resilient reachability. However, the slow convergence of prior works under network dynamics and uncertainties can hardly provide the dependability for tenants. Further, state-of-the-art load balance schemes are guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. In this paper, we propose vFab, a dependable virtualized fabric framework which can (1) quickly detect network failure in data plane, (2) explicitly select proper paths for all flows, and (3) converge to ideal bandwidth allocation at sub-millisecond. The core idea of vFab is to leverage the programmable data plane to build a fusion of an active edge (e.g., NIC) and an informative core (e.g., switch), where the core sends link status and tenant information to the edge via telemetry to help the latter make a timely and accurate decision on path selection and traffic admission. We fully implement vFab with commodity SmartNICs and programmable switches. Extensive evaluations show that vFab can keep bandwidth guarantee with high bandwidth utilization, low and bounded latency, and resilient reachability under various network scenarios with limited overhead. Application-level experiments show that vFab can improve QPS by$2.4\times $and cut tail latency by$10\times $compared to the alternatives. Kaihui Gao, Shuai Wang 0028, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Tao Sun 0010 |
IEEE/ACM Trans. Netw. | 9 |
| 2022 | Buffer-based End-to-end Request Event Monitoring in the Cloud
Kaihui Gao, Chen Sun 0005, Shuai Wang 0028, Dan Li 0001, Yu Zhou 0008, Hongqiang Harry Liu, Lingjun Zhu, Ming Zhang 0005 |
NSDI | 2 |
| 2022 | Predictable vFabric on informative data planeabstractIn multi-tenant data centers, each tenant desires reassuring predictability from the virtual network fabric - bandwidth guarantee, work conservation, and bounded tail latency. Achieving these goals simultaneously relies on rapid and precise traffic admission. However, the slow convergence (tens of milliseconds) of prior works can hardly satisfy the increasingly rigorous performance demand under dynamic traffic patterns. Further, state-of-the-art load balance schemes are all guarantee-agnostic and bring great risks on breaking bandwidth guarantee, which is overlooked in prior works. Shuai Wang 0028, Kaihui Gao, Kun Qian 0021, Dan Li 0001, Rui Miao 0001, Bo Li 0061, Yu Zhou 0008, Ennan Zhai, Chen Sun 0005, Binzhang Fu, Frank Kelly, Dennis Cai, Hongqiang Harry Liu, Ming Zhang 0005 |
SIGCOMM | 9 |
| 2022 | Achieving consistent low latency for wireless real-time communications with the shortest control loopabstractReal-time communication (RTC) applications like video conferencing or cloud gaming require consistent low latency to provide a seamless interactive experience. However, wireless networks including WiFi and cellular, albeit providing a satisfactory median latency, drastically degrade at the tail due to frequent and substantial wireless bandwidth fluctuations. We observe that the control loop for the sending rate of RTC applications is inflated when congestion happens at the wireless access point (AP), resulting in untimely rate adaption to wireless dynamics. Existing solutions, however, suffer from the inflated control loop and fail to quickly adapt to bandwidth fluctuations. In this paper, we propose Zhuge, a pure wireless AP based solution that reduces the control loop of RTC applications by separating congestion feedback from congested queues. We design a Fortune Teller to precisely estimate per-packet wireless latency upon its arrival at the wireless AP. To make Zhuge deployable at scale, we also design a Feedback Updater that translates the estimated latency to comprehensible feedback messages for various protocols and immediately delivers them back to senders for rate adaption. Trace-driven and real-world evaluation shows that Zhuge reduces the ratio of large tail latency and RTC performance degradation by 17% to 95%. Zili Meng, Yaning Guo, Chen Sun 0005, Bo Wang 0066, Justine Sherry, Hongqiang Harry Liu, Mingwei Xu 0001 |
SIGCOMM | 3 |
| 2022 | Firebolt: Finding Bugs in Programmable Data Plane Generators
Jiamin Cao, Yu Zhou 0008, Chen Sun 0005, Lin He 0004, Zhaowei Xi, Ying Liu 0024 |
USENIX ATC | 3 |
| 2022 | Newton: Intent-Driven Network Traffic MonitoringabstractNetwork monitoring systems are designed to fulfill operators’ intents and serve as essential tools to modern networks. As a result of rapidly increasing network bandwidth and scale nowadays, network monitors should satisfy on-demand network monitoring for continuously growing traffic volumes. However, existing monitoring systems either cannot satisfy flexible intents on demand or produce significant overheads. In this paper, we presentNewton, an intent-driven traffic monitor that is able to specify operators’ intents with traffic monitoring queries and conduct dynamic and scalable network-wide queries deployment.Newtonenables operators to customize and modify queries dynamically without interrupting the network workflow. Besides,Newtonproposes systematic optimizations at device level and network-wide level to reduce resource consumption while deploying queries.Newtoncan combine the resources across switches to deploy complex queries with high resilience to dynamic network status. Evaluations prove thatNewtonis of high flexibility, scalability, and resource efficiency, which demonstratesNewtonis promising to be deployed in large-scale programmable networks. Zhaowei Xi, Yu Zhou 0008, Kai Gao 0001, Chen Sun 0005, Jiamin Cao, Yangyang Wang 0001, Mingwei Xu 0001 |
IEEE/ACM Trans. Netw. | 5 |
| 2022 | CoFilter: High-Performance Switch-Accelerated Stateful Packet Filter for Bare-Metal ServersabstractAs one of the most critical cloud services, Bare-Metal Servers (BMS) introduce stringent performance requirements on data center networks (DCN). Stateful packet filter is an integral DCN component of ensuring connection security for BMS. However, the off-the-shelf stateful packet filters either are costly for cloud DCNs or introduce significant performance bottlenecks. In this article, we presentCoFilter, which leverages low-cost programmable switches to accelerate the stateful packet filter for BMS.CoFilteruses (1)stateful process partitionto enable complex stateful packet filtering logic on programmability-limited switching ASICs, (2)state compressionto track tens of millions of connections with constrained hardware memory, and (3)per-tenant packet rate limit and tenant-aware flow migrationto achieve efficient performance isolation among different tenants. Overall,CoFilterimplements a high-performance stateful packet filter via the co-design of programmable switching ASIC and CPU. We evaluateCoFilterunder various data center traffic traces with real-world flow distributions. The evaluation results show thatCoFilterremarkably outperforms NetFilter, i.e., forwarding packets at line rate (13x throughput of NetFilter), keeping packet delay within 1us, and freeing a significant quantity of CPU cores, with rather small memory usage, i.e., accommodating over$10^7$connections with only 16MB SRAM. Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Lin He 0004, Chen Sun 0005, Yangyang Wang 0001, Mingwei Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | Practically Deploying Heavyweight Adaptive Bitrate Algorithms With Teacher-Student LearningabstractMajor commercial client-side video players employ adaptive bitrate (ABR) algorithms to improve the user quality of experience (QoE). With the evolvement of ABR algorithms, increasingly complex methods such as neural networks have been adopted to pursue better performance. However, these complex methods are too heavyweight to be directly deployed in client devices with limited resources, such as mobile phones. Existing solutions suffer from a trade-off between algorithm performance and deployment overhead. To make the deployment of sophisticated ABR algorithms practical, we propose PiTree, a general, high-performance, and scalable framework that can faithfully convert sophisticated ABR algorithms into decision trees with teacher-student learning. In this way, network operators can train complex models offline and deploy converted lightweight decision trees online. We also present theoretical analysis on the conversion and provide two upper bounds of the prediction error during the conversion and the generalization loss after conversion. Evaluation on three representative ABR algorithms with both trace-driven emulation and real-world experiments demonstrates that PiTree could convert ABR algorithms into decision trees with <; 3% average performance degradation. Moreover, compared to original deployment solutions, PiTree could save considerable operating expenses for content providers. Zili Meng, Yaning Guo, Yixin Shen 0002, Chao Zhou 0003, Minhu Wang, Jia Zhang 0010, Mingwei Xu 0001, Chen Sun 0005, Hongxin Hu |
IEEE/ACM Trans. Netw. | 9 |
| 2021 | Octans: Optimal Placement of Service Function Chains in Many-Core SystemsabstractNetwork Function Virtualization (NFV) offers service delivery flexibility and reduces overall costs by running service function chains (SFCs) on commodity servers with many cores. Existing solutions for placing SFCs in one server treat all CPU cores as equal and allocate isolated CPU cores to network functions (NFs). However, advanced servers often adopt Non-Uniform Memory Access (NUMA) architecture to improve the scalability of many-core systems. CPU cores are grouped into nodes, incurring performance degradation due to cross-node memory access and intra-node resource contention. Our evaluation shows that randomly selecting cores to place NFs in an SFC could suffer from 39.2 percent lower throughput comparing to an optimal placement solution. In this article, we propose Octans, an NFV orchestrator to achieve maximum aggregate throughput of all SFCs in many-core systems. Octans first formulates the optimization problem as a Non-Linear Integer Programming (NLIP) Model. Then we identify the key factor for problem solving as evaluating the throughput drop of an NF caused by other NFs in the same SFC or different SFCs, i.e., performance drop index, and propose a formal and accurate prediction model based on system level performance metrics. Finally, we propose two online algorithms to quickly find near-optimal placement solutions for one-time and incremental deployment. Extensive evaluation on a prototype implementation shows that Octans significantly improves the aggregate throughput comparing to two state-of-the-art placement solutions by 27.1 ~ 45.2 percent for one-time deployment and by 20.9 ~ 38.1 percent for incremental deployment, with very low prediction errors. Moreover, Octans could quickly find a near-optimal placement solution with tiny optimality gap. Heng Yu 0005, Zhilong Zheng, Junxian Shen, Congcong Miao, Chen Sun 0005, Hongxin Hu, Jun Bi, Jilong Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | Newton: intent-driven network traffic monitoringabstractMonitoring network traffic based on operators' intents is essential to today's networks. As the bandwidth and size of networks increase steeply, monitoring systems shall fulfill the requirements of on-demand network monitoring for ever-growing traffic volumes. However, existing monitoring systems either cannot satisfy operators' intents on demand or introduce substantial monitoring overheads. In this paper, we present Newton, an intent-driven traffic monitor that enables specifying operators' intents with traffic monitoring queries and supports dynamic and scalable network-wide queries. Specifically, Newton 1) empowers operators to dynamically create, remove, and update on-data-plane queries without interrupting normal packet forwarding, 2) conducts systematic optimizations to achieve precise network traffic monitoring, and 3) executes network-wide queries with high resilience to dynamic network status. Evaluation results show that Newton improves the flexibility, scalability, and resource efficiency of traffic monitoring, demonstrating its great potential to be deployed in large-scale programmable networks. Yu Zhou 0008, Kai Gao 0001, Chen Sun 0005, Jiamin Cao, Yangyang Wang 0001, Mingwei Xu 0001 |
CoNEXT | 4 |
| 2020 | SmartChain: Enabling High-Performance Service Chain Partition between SmartNIC and CPUabstractSmart Network Interface Cards (SmartNICs) have been widely used to accelerate software-based network functions (NFs). However, from the scope of a service chain, a careless selection of NFs to offload onto SmartNIC could severely degrade the performance due to frequent communications between CPU and SmartNIC. In this paper, we present SmartChain, a high performance and efficient framework that achieves optimal partition of service chains between SmartNIC and CPU. SmartChain consists of two logical steps. First, SmartChain analyzes the suitability of elements in a chain to run on SmartNIC to exploit its high performance. Besides, SmartChain also ensures the dependencies between elements. Second, as our key novelty, SmartChain models the service chain latency and resource constraints, and solves the partition problem with 0-1 integer linear programming. We implement a SmartChain prototype based on Netronome SmartNIC. Evaluation results show that when used in real world cases, SmartChain could reduce the service chain latency by up to 87% with throughput maintained compared with strawman solutions. Shuhe Wang, Zili Meng, Chen Sun 0005, Minhu Wang, Mingwei Xu 0001, Jun Bi, Tong Yang 0003, Qun Huang 0001, Hongxin Hu |
ICC | 3 |
| 2020 | Martini: Bridging the Gap between Network Measurement and Control Using Switching ASICsabstractAdvanced network management systems, including network measurement and traffic control, rely on a remote controller to make control decisions. However, this approach incurs a long control loop of a few seconds to minutes. Even if we switch to switch-local controller, the latency is still tens of milliseconds and is unacceptable for many latency-sensitive tasks. In this paper, we propose Martini, a general framework that supports measurement-based timely control. The key idea is to perform measurement, control decision, and control entirely in the switch data plane. This could shorten the control loop of management tasks that require timely control based on only locally measured statistics in the switch. First, Martini introduces a set of primitives to describe management tasks. Next, Martini provides an innovative network-wide task placement mechanism to exploit resources of all switches to accommodate massive management tasks. Finally, Martini provides a code library and a compiler to support measurement and control on a state-of-the-art switching ASIC. Evaluation results show that Martini can effectively support a wide range of fine-timescale management tasks such as microburst detection and fast load balancing by reducing the control loop from seconds to nanoseconds. Shuhe Wang, Chen Sun 0005, Zili Meng, Minhu Wang, Jiamin Cao, Mingwei Xu 0001, Jun Bi, Qun Huang 0001, Masoud Moshref, Tong Yang 0003, Hongxin Hu, Gong Zhang 0001 |
ICNP | 2 |
| 2020 | Serpens: A High-Performance Serverless Platform for NFVabstractMany enterprises run Network Function Virtualization (NFV) services on public clouds to relieve management burdens and reduce costs. However, NFV operators still face the burden of choosing the right types of virtual machines (VMs) for various network functions (NFs), as well as the cost of renting VMs at a granularity of months or years while many VMs remain idle during valley hours. A recent computing model named serverless computing automatically executes user-defined functions on requests arrival, and charges users based on the number of processed requests. For NFV operators, serverless computing has the potential of completely relieving NF management burden and significantly reducing costs. Nevertheless, naively exploring existing serverless platforms for NFV introduces significant performance overheads in three aspects, including high remote state access latency, long NF launching time, and high packet delivery latency between NFs. To address these problems, we propose Serpens, a high-performance serverless platform for NFV. Firstly, Serpens designs a novel state management mechanism to support local state access. Secondly, Serpens proposes an efficient NF execution model to provide fast NF launching and avoid extra packet delivery. We have implemented a prototype of Serpens. Evaluation results demonstrate that Serpens could significantly improve performance for NFs and service function chains (SFCs) comparing to existing serverless platforms. Junxian Shen, Heng Yu 0005, Zhilong Zheng, Chen Sun 0005, Mingwei Xu 0001, Jilong Wang 0001 |
IWQoS | 4 |
| 2020 | Flow Event Telemetry on Programmable Data PlaneabstractNetwork performance anomalies (NPAs), e.g. long-tailed latency, bandwidth decline, etc., are increasingly crucial to cloud providers as applications are getting more sensitive to performance. The fundamental difficulty to quickly mitigate NPAs lies in the limitations of state-of-the-art network monitoring solutions --- coarse-grained counters, active probing, or packet telemetry either cannot provide enough insights on flows or incur too much overhead. This paper presents NetSeer, a flow event telemetry (FET) monitor which aims to discover and record all performance-critical data plane events, e.g. packet drops, congestion, path change, and packet pause. NetSeer is efficiently realized on the programmable data plane. It has a high coverage on flow events including inter-switch packet drop/corruption which is critical but also challenging to retrieve the original flow information, with novel intra- and inter-switch event detection algorithms running on data plane; NetSeer also achieves high scalability and accuracy with innovative designs of event aggregation, information compression, and message batching that mainly run on data plane, using switch CPU as complement. NetSeer has been implemented on commodity programmable switches and NICs. With real case studies and extensive experiments, we show NetSeer can reduce NPA mitigation time by 61%-99% with only 0.01% overhead of monitoring traffic. Yu Zhou 0008, Chen Sun 0005, Hongqiang Harry Liu, Rui Miao 0001, Bo Li 0061, Zhilong Zheng, Lingjun Zhu, Yongqing Xi, Dennis Cai, Ming Zhang 0005, Mingwei Xu 0001 |
SIGCOMM | 2 |
| 2020 | Lyra: A Cross-Platform Language and Compiler for Data Plane Programming on Heterogeneous ASICsabstractProgrammable data plane has been moving towards deployments in data centers as mainstream vendors of switching ASICs enable programmability in their newly launched products, such as Broadcom's Trident-4, Intel/Barefoot's Tofino, and Cisco's Silicon One. However, current data plane programs are written in low-level, chip-specific languages (e.g., P4 and NPL) and thus tightly coupled to the chip-specific architecture. As a result, it is arduous and error-prone to develop, maintain, and composite data plane programs in production networks. This paper presents Lyra, the first cross-platform, high-level language & compiler system that aids the programmers in programming data planes efficiently. Lyra offers a one-big-pipeline abstraction that allows programmers to use simple statements to express their intent, without laboriously taking care of the details in hardware; Lyra also proposes a set of synthesis and optimization techniques to automatically compile this "big-pipeline" program into multiple pieces of runnable chip-specific code that can be launched directly on the individual programmable switches of the target network. We built and evaluated Lyra. Lyra not only generates runnable real-world programs (in both P4 and NPL), but also uses up to 87.5% fewer hardware resources and up to 78% fewer lines of code than human-written programs. Ennan Zhai, Hongqiang Harry Liu, Rui Miao 0001, Yu Zhou 0008, Bingchuan Tian, Chen Sun 0005, Dennis Cai, Ming Zhang 0005, Minlan Yu |
SIGCOMM | 7 |
| 2019 | Bubble: Lightweight Core Sharing in NFVabstractMany researches have revealed the requirement of enabling multiple network functions (NFs) to share a CPU core in Network Function Virtualization (NFV) to support fine-grained NF models, efficient resource utilization, and chain consolidation. However, these works usually enable core sharing via kernel-level threads, which incurs significant performance degradation. In this paper, we present Bubble to enable lightweight core sharing in NFV. Bubble leverages user-level threads to eliminate the performance overhead introduced by kernel-level thread scheduling. Bubble is designed to satisfy unique requirements in NFV by providing accurate and low-overhead scheduling, in support of on- demand resource allocation, and accurate NF load measurement. Evaluations over a Bubble prototype implementation demonstrate that Bubble can improve the performance by 1.6Ã- to 6.2Ã- for co-located NFs and by 3.7Ã- to 68.8Ã- for a consolidated Service Function Chain (SFC) in a core against two state- of-the-art solutions. Haiping Wang 0002, Zhilong Zheng, Chen Sun 0005, Jun Bi |
GLOBECOM | 3 |
| 2019 | Buffet: Enabling Multi-Tenant Network FunctionsabstractMany enterprises outsource traffic processing to third- party Network Function (NF) service providers to relieve management burden and reduce cost. NF providers have to process packets from multiple tenants simultaneously. However, most existing software based NFs are designed for one single tenant without internal state isolation mechanisms. These NFs cannot be securely shared across multiple tenants. Existing solutions that support multitenancy are either inefficient or ad-hoc for specific NFs. In this paper, we propose Buffet, a general and efficient framework that enables multitenancy for a wide range of NFs. First, Buffet introduces a general programming abstraction for various NFs to relieve NF developers from considering isolation details. Second, Buffet proposes dynamic tenant-level affinity to achieve high performance and resource efficiency. Finally, Buffet exploits SmartNIC offloading to eliminate host CPU overhead. We have implemented a prototype of Buffet. Evaluation results demonstrate that Buffet can effectively enable multitenancy for a wide range of NFs with high performance and resource efficiency. Heng Yu 0005, Junxian Shen, Chen Sun 0005, Zhilong Zheng, Jilong Wang 0001 |
GLOBECOM | 3 |
| 2019 | CoFilter: A High-Performance Switch-Accelerated Stateful Packet Filter for Bare-Metal ServersabstractAs one of the most critical cloud services, Bare-metal Servers introduce stringent performance requirements on data center networks (DCN). Stateful packet filter is an integral DCN component of ensuring connection security for bare-metal servers. However, the off-the-shelf hardware-based and software-based stateful packet filters either are prohibitively costly for cloud DCNs or introduce significant performance bottlenecks. In this paper, we present CoFilter, which employs cheap programmable switches to accelerate the stateful packet filter for bare-metal servers. CoFilter consists of two key designs. First, to support complex stateful packet filtering logic in programmability-limited switching ASICs, CoFilter partitions the stateful packet filtering logic between programmable ASICs and switch CPU. Most packets are directly processed in switching ASICs to achieve high performance, while only a small number of packets go to switch CPU for connection tracking. Second, to track massive connections with constrained hardware memory, CoFilter employs hash to compress connection states and provides an efficient settlement for hash collisions. We build a prototype of CoFilter and evaluate it on the Tofino switch under various data center traffic traces with real-world flow distribution. The evaluation shows that CoFilter largely outperforms NetFilter, i.e., forwarding packets at line rate (13x throughput of NetFilter), keeping packet delay at 1us, and freeing a significant quantity of CPU cores. Furthermore, CoFilter presents great scalability and accommodates over ten million connections with only 16MB SRAM. Jiamin Cao, Ying Liu 0024, Yu Zhou 0008, Chen Sun 0005, Yangyang Wang 0001, Jun Bi |
ICCCN | 4 |
| 2019 | Octans: Optimal Placement of Service Function Chains in Many-Core SystemsabstractNetwork Function Virtualization (NFV) has the potential to offer service delivery flexibility and reduce overall costs by running service function chains (SFCs) on commodity servers with many cores. Existing solutions for placing SFCs in one server treat all CPU cores as equal and allocate isolated CPU cores to different network functions (NFs). However, advanced servers often adopt Non-Uniform Memory Access (NUMA) architecture to improve the scalability of many-core systems. CPU cores are grouped into nodes, incurring performance bottleneck due to cross-node memory access and intra-node resource contention. Our evaluation shows that randomly selecting cores to place NFs in an SFC could suffer from 39.2% lower throughput comparing to an optimal placement solution. In this paper, we propose Octans, an NFV orchestrator to achieve maximum aggregate throughput of all SFCs in many-core systems. Octans first formulates the optimization problem as a Non-Linear Integer Programming (NLIP) model. Then we identify the key factor for problem solving as evaluating the throughput drop of an NF caused by other NFs in the same SFC or different SFCs, i.e. performance drop index, and propose a formal and precise prediction model based on system level performance metrics. Finally, we propose an efficient heuristic algorithm to quickly find near-optimal placement solutions. We have implemented a prototype of Octans. Extensive evaluation shows that Octans significantly improves the aggregate throughput comparing to two state-of the-art placement mechanisms by 26.7%~51.8%, with very low prediction errors of SFC performance (an average deviation of 2.6%). Moreover, Octans could quickly find a near-optimal placement solution with tiny optimality gap (1.2%~3.5%). Zhilong Zheng, Jun Bi, Heng Yu 0005, Haiping Wang 0002, Chen Sun 0005, Hongxin Hu |
INFOCOM | 5 |
| 2019 | P4Tester: efficient runtime rule fault detection for programmable data planesabstractP4 and programmable data planes bring significant flexibility to network operation but are inevitably prone to various faults. Some faults, like P4 program bugs, can be verified statically, while some faults, like runtime rule faults, only happen to running network devices, and they are hardly possible to handle before deployment. Existing network testing systems can troubleshoot runtime rule faults via injecting probes, but are insufficient for programmable data planes due to large overheads or limited fault coverage. In this paper, we propose P4Tester, a new network testing system for troubleshooting runtime rule faults on programmable data planes. First, P4Tester proposes a new intermediate representation based on Binary Decision Diagram, which enables efficient probe generation for various P4-defined data plane functions. Second, P4Tester offers a new probe model that uses source routing to forward probes. This probe model largely reduces rule fault detection overheads, i.e. requiring only one server to generate probes for large networks and minimizing the number of probes. Moreover, this probe model can test all table rules in a network, achieving full fault coverage. Evaluation based on real-world data sets indicates that P4Tester can efficiently check all rules in programmable data planes, generate 59% fewer probes than ATPG and Pronto, be faster than ATPG by two orders of magnitude, and troubleshoot multiple rule faults within one second on BMv2 and Tofino. Yu Zhou 0008, Jun Bi, Yunsenxiao Lin, Yangyang Wang 0001, Zhaowei Xi, Jiamin Cao, Chen Sun 0005 |
IWQoS | 8 |
| 2019 | PiTree: Practical Implementation of ABR Algorithms Using Decision TreesabstractMajor commercial client-side video players employ adaptive bitrate (ABR) algorithms to improve user quality of experience (QoE). With the evolvement of ABR algorithms, increasingly complex methods such as neural networks have been adopted to pursue better performance. However, these complex methods are too heavyweight to be directly implemented in client devices, especially mobile phones with very limited resources. Existing solutions suffer from a trade-off between algorithm performance and deployment overhead. To make the implementation of sophisticated ABR algorithms practical, we propose PiTree, a general, high-performance and scalable framework that can faithfully convert sophisticated ABR algorithms into lightweight decision trees to reduce deployment overhead. We also provide a theoretical upper bound on the optimization loss during the conversion. Evaluation results on three representative ABR algorithms demonstrate that PiTree could faithfully convert ABR algorithms into decision trees with <3% average performance degradation. Moreover, comparing to original implementation solutions, PiTree could save operating expenses for large content providers. Zili Meng, Yaning Guo, Chen Sun 0005, Hongxin Hu, Mingwei Xu 0001 |
ACM Multimedia | 4 |
| 2019 | MicroNF: An Efficient Framework for Enabling Modularized Service Chains in NFVabstractThe modularization of service function chains (SFCs) in network function virtualization (NFV) could introduce significant performance overhead and resource efficiency degradation due to introducing frequent packet transfer and consuming much more hardware resources. In response, we exploit the reusability, lightweightness, and individual scalability features of elements in modularized SFCs (MSFCs) and propose MicroNF, an efficient framework for MSFC in NFV. MicroNF addresses the performance overhead and resource efficiency problems in three ways. First, MicroNF graph constructor reuses the processing results of elements from different NFs and reconstructs the MSFC after modularization to shorten the chain latency. Second, optimized placer pays attention to the problem of which elements to consolidate and provides a performance-aware placement algorithm to place MSFCs compactly and optimize the global packet transfer cost. Third, MicroNF individual scaler innovatively introduces a push-aside scaling up strategy to avoid degrading performance and taking up new CPU cores. To support MSFC reusing and consolidation, MicroNF also designs a high-performance infrastructure to efficiently forwarding packets with consistency ensured and to automatically scheduling elements with fairness ensured when the elements are consolidated on the CPU core. Our evaluation results show that MicroNF achieves significant performance improvement and efficient resource utilization on several metrics. Zili Meng, Jun Bi, Haiping Wang 0002, Chen Sun 0005, Hongxin Hu |
IEEE J. Sel. Areas Commun. | 4 |
| 2018 | GEN: A GPU-Accelerated Elastic Framework for NFVabstractNetwork Function Virtualization (NFV) has the potential to enhance service delivery flexibility and reduce overall costs by provisioning software-based service function chains (SFCs) on commodity hardware. However, we observe that existing CPU-based SFC solutions cannot achieve both high performance and high elasticity simultaneously. To address such a critical challenge, we seek beyond CPU and exploit the capability of Graphics Processing Unit (GPU) to support NFV. We propose GEN, a GPU-based high performance and elastic framework for NFV. As opposed to pipeline-based SFCs in existing GPU-based NFV systems, GEN proposes to support RTC-based SFCs to improve processing performance. Meanwhile, GEN offers great elasticity of network function (NF) scaling up and down by allocating a different number of fine-grained GPU threads to an NF during runtime. We have implemented a prototype of GEN. Preliminary evaluation results demonstrate that GEN improves performance with RTC-based SFCs, and supports adaptive, precise, and fast NF scaling for NFV. Zhilong Zheng, Jun Bi, Chen Sun 0005, Heng Yu 0005, Hongxin Hu, Zili Meng, Shuhe Wang, Kai Gao 0001 |
APNet | 3 |
| 2018 | CoCo: Compact and Optimized Consolidation of Modularized Service Function Chains in NFVabstractThe modularization of Service Function Chains (SFCs) in Network Function Virtualization (NFV) could introduce significant performance overhead and resource efficiency degradation due to introducing frequent packet transfer and consuming much more hardware resources. In response, we exploit the lightweight and individually scalable features of elements in Modularized SFCs (MSFCs) and propose CoCo, a compact and optimized consolidation framework for MSFC in NFV. CoCo addresses the above problems in two ways. First, CoCo Optimized Placer pays attention to the problem of which elements to consolidate and provides a performance-aware placement algorithm to place MSFCs compactly and optimize the global packet transfer cost. Second, CoCo Individual Scaler innovatively introduces a push-aside scaling up strategy to avoid degrading performance and taking up new CPU cores. To support MSFC consolidation, CoCo also provides an automatic runtime scheduler to ensure fairness when elements are consolidated on CPU core. Our evaluation results show that CoCo achieves significant performance improvement and efficient resource utilization. Zili Meng, Jun Bi, Haiping Wang 0002, Chen Sun 0005, Hongxin Hu |
ICC | 4 |
| 2018 | Grus: Enabling Latency SLOs for GPU-Accelerated NFV SystemsabstractGraphics Processing Unit (GPU) has been recently exploited as a hardware accelerator to improve the performance of Network Function Virtualization (NFV). However, GPU-accelerated NFV systems suffer from significant latency variation when multiple network functions (NFs) are co-located in the same machine, which prevents operators from supporting latency Service Level Objectives (SLOs). Existing research efforts to address this problem can only guarantee a limited number of SLOs with very low resource utilization efficiency. In this paper, we present the Grus framework to support latency SLOs in GPU-accelerated NFV systems. Grus thoroughly analyzes the sources of latency variation and proposes three design principles: (1) dynamic batch size setting is needed to bound packet batching latency in CPU; (2) a reordering mechanism for data transfer over PCI-E is required to guarantee the stalling time; and (3) maximizing concurrency in GPU is necessary to avoid NF execution waiting time. Guided by the principles, Grus consists of two logical layers including an infrastructure layer and a scheduling layer. The infrastructure layer is equipped with an in-CPU Reorder-able Worker Pool that could adjust batching size and packet transfer order, and in-GPU Controllable Concurrent Executors to provide maximized concurrency. The scheduling layer runs a heuristic algorithm to perform accurate and fast scheduling to guarantee SLOs based on our prediction models. We have implemented a prototype of Grus. Extensive evaluations demonstrate that Grus can significantly reduce latency variation and satisfy 4.5 × more SLO terms than state-of-the-art solutions. Zhilong Zheng, Jun Bi, Haiping Wang 0002, Chen Sun 0005, Heng Yu 0005, Hongxin Hu, Kai Gao 0001 |
ICNP | 4 |
| 2018 | OFM: Optimized Flow Migration for NFV Elasticity ControlabstractNetwork Function Virtualization (NFV) together with Software Defined Networking (SDN) offers the potential for enhancing service delivery flexibility and reducing overall costs. Based on the capability of dynamic creation and destruction of network function (NF) instances, NFV provides great elasticity in NF control, such as NF scaling out, scaling in, load balancing, etc. To realize NFV elasticity control, network traffic flows need to be redistributed across NF instances. However, deciding which flows are suitable for migration is a critical problem for efficient NFV elasticity control. In this paper, we propose to build an innovative flow migration controller, OFM Controller, to achieve optimized flow migration for NFV elasticity control. We identify the trigger conditions and control goals for different situations, and carefully design models and algorithms to address three major challenges including buffer overflow avoidance, migration cost calculation, and effective flow selection for migration. We implement the OFM Controller on top of NFV and SDN environments. Our evaluation results show that OFM Controller is efficient to support optimized flow migration in NFV elasticity control. Chen Sun 0005, Jun Bi, Zili Meng, Hongxin Hu |
IWQoS | 1 |
| 2018 | Enabling NFV Elasticity Control With Optimized Flow MigrationabstractNetwork function virtualization (NFV) together with software defined networking (SDN) offers the potential for enhancing service delivery flexibility and reducing overall costs. Based on the capability of dynamic creation and destruction of network function (NF) instances, NFV provides great elasticity in NF control, such as NF scaling out, scaling in, and load balancing. To realize NFV elasticity control, network traffic flows need to be redistributed across NF instances. However, deciding which flows are suitable for migration is a critical problem for efficient NFV elasticity control. In this paper, we propose to build an innovative flow migration controller, OFM controller, to achieve optimized flow migration for NFV elasticity control. We identify the trigger conditions and control goals for different situations, and carefully design models and algorithms to address three major challenges including buffer overflow avoidance, migration cost calculation, and effective flow selection for migration. We implement the OFM controller on top of NFV and SDN environments. Our evaluation results show that OFM controller is efficient to support optimized flow migration in NFV elasticity control. Chen Sun 0005, Jun Bi, Zili Meng, Tong Yang 0003, Hongxin Hu |
IEEE J. Sel. Areas Commun. | 1 |
| 2017 | NFP: Enabling Network Function Parallelism in NFVabstractSoftware-based sequential service chains in Network Function Virtualization (NFV) could introduce significant performance overhead. Current acceleration efforts for NFV mainly target on optimizing each component of the sequential service chain. However, based on the statistics from real world enterprise networks, we observe that 53.8% network function (NF) pairs can work in parallel. In particular, 41.5% NF pairs can be parallelized without causing extra resource overhead. In this paper, we present NFP, a high performance framework, that innovatively enables network function parallelism to improve NFV performance. NFP consists of three logical components. First, NFP provides a policy specification scheme for operators to intuitively describe sequential or parallel NF chaining intents. Second, NFP orchestrator intelligently identifies NF dependency and automatically compiles the policies into high performance service graphs. Third, NFP infrastructure performs light-weight packet copying, distributed parallel packet delivery, and load-balanced merging of packet copies to support NF parallelism. We implement an NFP prototype based on DPDK in Linux containers. Our evaluation results show that NFP achieves significant latency reduction for real world service chains. Chen Sun 0005, Jun Bi, Zhilong Zheng, Heng Yu 0005, Hongxin Hu |
SIGCOMM | 1 |
| 2017 | HYPER: A Hybrid High-Performance Framework for Network Function VirtualizationabstractNetwork function virtualization (NFV) offers the potential for both enhancing service delivery flexibility and reducing overall costs by virtualizing network functions that are traditionally implemented in dedicated hardware. However, the flexibility of NFV comes with considerable compromises since virtual machine carried functions could introduce significant performance overhead. In this paper, we present a novel high-performance framework called HYPER, which combines programmable hardware infrastructure and traditional software infrastructure in NFV to achieve both high performance and flexibility for supporting virtualized network functions (VNFs). In HYPER, we design a mediator layer to hide underlying infrastructure heterogeneity from the NFV orchestrator to simplify VNF management. In addition, we design a SLA-aware service chaining algorithm in HYPER to leverage the benefits of the hybrid infrastructure to fulfill both functional and performance requirements from service subscribers (or tenants). To optimize resource utilization efficiency, we also introduce a performance-aware VNF placement algorithm in HYPER, which accommodates both resource and performance requirements in placing VNFs. We implement HYPER in a testbed based on OpenStack and ONetCard. Experimental results show that HYPER reduces the forwarding latency of a service chain by 40% to 67% compared with data plane development kit -based implementation, while maintaining the flexibility of VNF management. Chen Sun 0005, Jun Bi, Zhilong Zheng, Hongxin Hu |
IEEE J. Sel. Areas Commun. | 1 |
| 2017 | SDPA: Toward a Stateful Data Plane in Software-Defined NetworkingabstractAs the prevailing technique of software-defined networking (SDN), open flow introduces significant programmability, granularity, and flexibility for many network applications to effectively manage and process network flows. However, open flow only provides a simple “match-action” paradigm and lacks the functionality of stateful forwarding for the SDN data plane, which limits its ability to support advanced network applications. Heavily relying on SDN controllers for all state maintenance incurs both scalability and performance issues. In this paper, we propose a novel stateful data plane architecture (SDPA) for the SDN data plane. A co-processing unit, forwarding processor (FP), is designed for SDN switches to manage state information through new instructions and state tables. We design and implement an extended open flow protocol to support the communication between the controller and FP. To demonstrate the practicality and feasibility of our approach, we implement both software and hardware prototypes of SDPA switches, and develop a sample network function chain with stateful firewall, domain name system (DNS) reflection defense, and heavy hitter detection applications in one SDPA-based switch. Experimental results show that the SDPA architecture can effectively improve the forwarding efficiency with manageable processing overhead for those applications that need stateful forwarding in SDN-based networks. Chen Sun 0005, Jun Bi, Haoxian Chen 0001, Hongxin Hu, Zhilong Zheng, Shuyong Zhu, Chenghui Wu |
IEEE/ACM Trans. Netw. | 1 |
| 2016 | NeSMA: Enabling network-level state-aware applications in SDNabstractAs the de facto data plane technique of Software-Defined Networking (SDN), OpenFlow introduces significant programmability to enable innovative network applications. However, the simple OpenFlow data plane only maintains flow-level counters and lacks an efficient mechanism to manage network-level states, which limits its support for advanced state-aware applications. Regularly pulling whole state information from the data plane to the controller might incur untimely response to important network-level states such as CPU exhaustion, switch overload, etc and cause unnecessary traffic. To address above challenges, we introduce a novel Network-level State Management Architecture (NeSMA) to efficiently support advanced network-level state-aware applications by exploiting the opportunity of SDN central control. The data plane could be configured to check state regularly and report to the controller when triggered by state transitions. We design both sequential and parallel composition methods to deal with complex network-level states in NeSMA. To demonstrate the feasibility of our approach, we implement a software prototype of NeSMA, based on which we develop a data-center flow scheduling application. Experimental results show that NeSMA can process network-level states with low network resource consumption and high scalability without compromising packet forwarding efficiency. Chen Sun 0005, Jun Bi, Hongxin Hu, Zhilong Zheng |
ICNP | 1 |
| 2016 | SLA-NFV: an SLA-aware High Performance Framework for Network Function VirtualizationabstractWe propose SLA-NFV, a Service Level Agreement (SLA) aware framework, for building high-performance NFV, focusing on fulfilling SLAs of service subscribers (or tenants). SLA-NFV leverages a hybrid infrastructure with both software and programmable hardware to enhance NFV’s capability with respect to various SLAs. Evaluations show that a hybrid service chain could reduce latency by up to 60% compared with a pure soft- ware service chain. Chen Sun 0005, Jun Bi, Zhilong Zheng, Hongxin Hu |
SIGCOMM | 1 |
| 2015 | SDPA: Enhancing Stateful Forwarding for Software-Defined NetworkingabstractAs the prevailing technique of Software-Defined Networking (SDN), OpenFlow introduces significant programmability, granularity and flexibility for many network applications to effectively manage and process network flows. However, OpenFlow only provides a simple "match-action" paradigm and lacks the function of stateful forwarding for SDN data plane, which limits it to support advanced network applications. Heavily relying on SDN controllers for all state maintenance incurs both scalability and performance issues. In this paper, we propose a novel Stateful Data Plane Architecture (SDPA) for SDN data plane. A co-processing unit, Forwarding Processor (FP), is designed for SDN switches to manage state information through new instructions and state tables. We design and implement an extended OpenFlow protocol to implement the communication between the controller and FP. To demonstrate the practicality and feasibility of our approach, we implement both software and hardware prototypes of SDPA switches, and develop a sample network function chain with stateful firewall, DNS reflection attack defense and NAT applications in one SDPA-based switch. Experimental results show that the SDPA architecture can effectively improve the forwarding efficiency with manageable processing overhead for those applications that need stateful forwarding in SDN-based networks. Shuyong Zhu, Jun Bi, Chen Sun 0005, Chenhui Wu, Hongxin Hu |
ICNP | 3 |