VLDB 2026 Research / reviewers in the wild / expert
Zhigang Sun 0002
dblp:20/3089-2 · also Zhi-gang Sun 0002, ZhiGang Sun 0002
· DBLP profile ↗
36ranked-venue papers
0as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 19 · 5 since 2021Systems, architecture and hardware · 9 · 4 since 2021Security and privacy · 2Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PDE-TSN: Enable TSN Autonomous Self-healing under Link Faults
Wenwen Fu, Xuyan Jiang, Wei Quan 0004, Tao Li 0008, Zhigang Sun 0002 |
SECON | 8 |
| 2026 | KaleidoScope: A Co-Processor for Neural-Network-Driven Intelligent Data Plane
Dong Wen 0004, Zhongpei Liu, Tong Yang 0003, Tianyun Li, Yanshu Wang, Tao Li 0008, Zhuochen Fan, Qing Li 0006, Zhigang Sun 0002 |
IEEE Trans. Computers | 9 |
| 2026 | Pao-Ding: Accelerating Cross-Edge Video Analytics via Automated CNN Model Partitioning
Guanping Liang, Biao Han 0003, Ruidong Li 0001, Xueqiang Han, Zhigang Sun 0002 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | FooDog: Empower TSN for Efficient PolicingabstractTime-Sensitive Networking (TSN) is an emerging real-time Ethernet technology that provides deterministic communication for time-sensitive (TS) traffic. At its core, TSN utilizes Per-Stream Filtering and Policing (PSFP) gates to mitigate the disruption of unavoidable frame drift. However, as first identified in this work, the naive PSFP gate design results in heavy memory usage, which hinders normal switching functions. This work proposes an efficient PSFP gate design called FooDog. FooDog employs a two-stage structure and a dual-engine policing mechanism to realize memory-efficient, logic-compact, and fast policing while maintaining minimal latency and jitter for TS traffic. Results on FPGA prototypes show that FooDog consumes only hundreds of kilobits of memory, reducing on-chip memory overheads by more than 90% compared to the unoptimized PSFP gate design. Additionally, it maintains end-to-end latency in the microsecond range and jitter below 150 nanoseconds under abnormal traffic conditions, comparable to typical TSN performance without anomalies. Xuyan Jiang, Xiangrui Yang 0002, Tongqing Zhou, Wenfei Wu, Wenwen Fu, Wei Quan 0004, Yingwen Chen 0001, Yihao Jiao, Zhigang Sun 0002 |
IEEE Trans. Netw. | 9 |
| 2025 | A Performance-Balanced Scheduling Algorithm for Diverse Real-World TSN ScenariosabstractTime-Sensitive Networking (TSN) achieves low-delay and low-jitter traffic transmission through different traffic scheduling mechanisms. However, despite numerous algorithms developed based on these mechanisms, most fail to concurrently support multipath, hybrid, and multicast traffic, which are prevalent in real-world scenarios. Moreover, balancing performance metrics such as success rate, bandwidth utilization, and computation overhead remains challenging for these algorithms, significantly limiting their application in diverse TSN scenarios. To solve this problem, this paper proposes a universal ultra-low-delay and zero-jitter traffic scheduling model. Based on this model, this paper further designs a performance-balanced algorithm. The algorithm improves traffic scheduling success rate through joint routing and scheduling, increases network bandwidth utilization through hybrid traffic scheduling, and achieves low computation overhead through policy-based searching. Finally, extensive experiments demonstrate that the algorithm effectively balances performance metrics across diverse real-world scenarios. It achieves high scheduling success rate under real-world traffic loads ($\gt $20% improvement over non-joint routing), increased bandwidth utilization in the presence of hybrid traffic (18.3% enhancement over non-hybrid traffic scheduling), and low computation overhead ($\lt $2 minutes). Xuyan Jiang, Rulin Liu, Tao Li 0008, Wei Quan 0004, Zhigang Sun 0002 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2024 | Node Bundle Scheduling: An Ultra-low Latency Traffic Scheduling Algorithm for TAS-Based Time-Sensitive Networks
Xuyan Jiang, Wei Quan 0004, Rulin Liu, Zhigang Sun 0002 |
Euro-Par (1) | 5 |
| 2024 | Hebo: FPGA-based Transfer Time Planning for Volatile Traffic in TSNabstractTime-Sensitive Networking (TSN) is an advanced technology designed for real-time Ethernet communications, providing extremely low latency, minimal jitter, and lossless data transfer for time-sensitive critical traffic. Despite its benefits, TSN faces challenges with volatile traffic, where the time between frames constantly changes, leading to potential network performance issues. To tackle this issue, this paper introduces Hebo, a novel solution designed for zero frame loss and minimal latency of volatile traffic. Hebo employs a centralized controller, built with field-programmable gate array (FPGA), to dynamically plan the timing of volatile traffic in real-time. This approach ensures real-time data transmission across the network by efficiently allocating network resources. Our real-world and simulation experiments show that Hebo could significantly improve network performance, achieving less than 100 microseconds in end-to-end delay and eliminating nearly all frame loss (reducing it from approximately 80% to zero) in industrial automation scenarios. Xuyan Jiang, Zitong Wang 0002, Xiangrui Yang 0002, Yihao Jiao, Tianci Yu, Wenwen Fu, Yinhan Sun, Zhigang Sun 0002 |
IWQoS | 9 |
| 2023 | Fenglin-I: An Open-Source Time-Sensitive Networking Chip Enabling Agile CustomizationabstractTime-Sensitive Networking (TSN) technology is experiencing diverse application requirements and forming a complicated standard system. It is extremely difficult to design a one-fits-all chip for all TSN applications. Therefore, application-driven TSN chip customization is inevitable. Generally, chip customization starts from a “clean-slate”. For complicated ASIC chips, that results in significant development overhead. Inspired by RISC-V chips, an open-source template will significantly reduce the customization complexity. Along this road, we propose an open-source TSN chip named Fenglin-I. Fenglin-I includes a high-level abstraction to build a relationship between application requirements and chip implementation, source code of a real chip named FastTSN to provide reference code for chip implementation, and software tools to facilitate chip verification. Based on Fenglin-I, we further propose a TSN chip customization method that provides step-by-step guidance about customizing TSN chips agilely. To verify the effectiveness of Fenglin-I and the proposed customization method, we use FPGA arrays to prototype and verify FastTSN. The results show that FastTSN achieves microsecond-level transmission jitter for unicast and multicast time-critical traffic. Additionally, we demonstrate two domain-specific TSN chip customization cases in which the customized chips reuse at least 84$\%$of FastTSN code while meeting their requirements. Wenwen Fu, Wei Quan 0004, Jinli Yan, Zhigang Sun 0002 |
IEEE Trans. Computers | 4 |
| 2022 | TASP: Enabling Time-Triggered Task Scheduling in TSN-Based Mixed-Criticality SystemsabstractDistributed mixed-criticality system (DMCS) has been widely used in various critical domains such as self-driving cars and space crafts. To guarantee the end-to-end QoS (i.e., deadline/jitter requirements) of sensing-controlling-actuating control loops (CL) applications, DMCS adopts Time-Sensitive Networking (TSN), an emerging real-time Ethernet technology, for communication between end systems (ES). TSN provides a synchronized global clock and guarantees bounded delay for time-critical traffic in CLs, making it possible for DMCS to collaboratively schedule the computation (on ES) and communication (in TSN) to meet the Quality of Service (QoS) requirement. However, as modern DMCS tends to use fully-fledged Linux distributions (rather than a custom real-time OS) on ES to enjoy Linux’s mature ecosystem, it is challenging for DMCS to realize TSN-based QoS guarantees because the event-triggered scheduling of Linux on ES is incompatible with TSN.This paper proposes TAsk Scheduling Puppeteer (TASP), a mechanism that schedules CL tasks based on TSN without modifying the Linux OS. The key idea of TASP is to manipulate task scheduling by controlling the timing of CL packet submissions at the interface between TSN and ES. Specifically, TASP extracts two parameters: Fore Guardband (ForeGB) and Back Guardband (BackGB). During the ForeGB period before a CL packet’s submission, TASP forbids any packets’ submission; while during the BackGB period after a CL packet’s submission, TASP forbids any other CL packets’ submission. ForeGB and BackGB can ensure that there is at most one schedulable task on the ES at any time, and thus Linux has no choice but to schedule the only task, making the Linux scheduler a puppet. We have deployed TASP and evaluated it in real-world TSN switches based on an open-source TSN project, OpenTSN. The TASP-enabled ES can achieve task scheduling precisely based on TSN’s global clock, which outperforms the original ES by reducing end-to-end jitter from milliseconds to microseconds. Xuyan Jiang, Yiming Zhang 0003, Wenwen Fu, Xiangrui Yang 0002, Yinhan Sun, Zhigang Sun 0002 |
IWQoS | 6 |
| 2020 | TSN-Builder: Enabling Rapid Customization of Resource-Efficient Switches for Time-Sensitive NetworkingabstractTime-Sensitive Networking (TSN) emerges as a promising technique empowering deterministic forwarding on standard Ethernet without sacrificing compatibility. There are some commercial off-the-shelf (COTS) switches that support TSN recently. However, the resource partitioning on these switches is normally inefficient for the on-chip memory resource in many specific application scenarios. We observe that the critical requirements (e.g., topology, flow features) of these scenarios are pre-determined. Thus, developing a TSN switch in a Top-down approach is feasible and urgently needed.In this paper, we propose TSN-Builder, a template-based developing model for customizing resource-efficient TSN switches rapidly with targeted application-dependent requirements. TSN-Builder decomposes the integrated TSN switching function into multiple function templates. With a fine-grained resource abstraction, TSN-Builder provides platform-independent customization interfaces for developers to customize the resource parameters. We prototype TSN switches on FPGA to evaluate the resource consumption and performance under different application scenarios. Experimental results show that TSN-Builder reduces the on-chip memory by up to 80.53% under the same Quality-of-Service, compared to the resource configuration in the COTS switch. Jinli Yan, Wei Quan 0004, Xiangrui Yang 0002, Wenwen Fu, Zhigang Sun 0002 |
DAC | 7 |
| 2020 | Injection Time Planning: Making CQF Practical in Time-Sensitive NetworkingabstractTime-Aware Shaper (TAS) is a core mechanism to guarantee the deterministic transmission for periodic time-sensitive flows in Time-Sensitive Networking (TSN). The generic TAS requires complex configurations for the Gate Control List (GCL) attached to each queue in a switch. To simplify the design of a TSN switch, a Ping-Pong queue-based model named Cyclic Queuing and Forwarding (CQF) was proposed in IEEE 802.1 Qch by assigning fixed configurations to TAS. However, IEEE 802.1 Qch only defines the queue model and workflow of CQF. A global planning mechanism which maps the time-sensitive flows onto the underlying resources both temporally and spatially is urgently needed to make CQF practical.In this paper, we propose an Injection Time Planning (ITP) mechanism to optimize the network throughput of time-sensitive flows based on the observation that the start time when the packets are injected into the network has an important influence on the utilization of CQF queue resources. ITP provides a global temporal and spatial resource abstraction to make the implementation details transparent to algorithm designers. Based on our ITP mechanism, a novel heuristic algorithm named Tabu-ITP with domain-specific optimizing strategies is designed and evaluated under three typical network topologies in industrial control scenarios. Compared with the Naive algorithm without using ITP mechanism, experimental results demonstrate that Tabu-ITP improves the mapped flow number by 10x and the resource utilization by 65%. Jinli Yan, Wei Quan 0004, Xuyan Jiang, Zhigang Sun 0002 |
INFOCOM | 4 |
| 2020 | A Hierarchical Model of Control Logic for Simplifying Complex Networks Protocol Design
Wei Quan 0004, Jinli Yan, Zhigang Sun 0002 |
NPC | 5 |
| 2019 | STRIDE: Single-Trip-Time Based Reliable Data Transport Protocol for the Reconfigurable CloudabstractIn a recent development, reconfigurable clouds become a viable solution to overcome practical problems in clouds, such as scalability, delay, etc., by offloading computation tasks to reconfigurable hardware, FPGA. Several existing techniques, such as TCP/IP Offload Engine (TOE) and Lightweight Transport Layer (LTL), are still hard to be implemented in real-world deployment due to large overhead or stringent dependency of the underlying network. In this paper, we propose STRIDE, a novel inter-FPGA data communication protocol, to provide reliable end-to-end communication which addresses practical problems in deployment. In our design, STRIDE leverages FPGA's abilities through programming, such as precise timestamping, to make more accurate measurement on end-to-end delay and deliver more precise control in managing traffic in cloud. We implement STRIDE on a FPGA-based network experimental platform and demonstrate that STRIDE reduces various hardware resources consumption by 36% to 49% compared to TOE. Additionally, it also improves flow completion time in comparison to TOE by 2.2X. We further demonstrate STRIDE outperforms DCTCP and TCP-Vegas in OMNET simulator by up to 1.8X and 2.3X on average and 99th percentile respectively in large scale setting. Wenwen Fu, Tao Li 0008, Jialun Yang, Junnan Li 0002, Zhigang Sun 0002 |
ICC | 5 |
| 2019 | A Heterogeneous Parallel Packet Processing Architecture for NFV AccelerationabstractNetwork function virtualization (NFV) offers a new way to design, deploy and manage networking services. It is of vital importance to exploit heterogeneous parallelism between hardware and software, in order to improve virtulization performance and quality of virtualized network services. In this poster, we propose a novel heterogeneous parallel architecture that highly exploits the parallelism inside packet processing, and implementation efficacy with hardware processing engines and software threads. We present two packet processing pipelines with three implemented VNF instances to better demonstrate the efficiency of heterogeneous parallelism in accelerating NFV. We show the performance of our proposed architecture with various virtualized requirements and traffics in a well-deployed network environment. Experimental results reveal that it can achieve accelerated NFV performance, as well as provide a wide class of VNFs to improve the quality of virtualized network services. Jinshu Su, Biao Han 0003, Gaofeng Lv, Tao Li 0008, Zhigang Sun 0002 |
ICNP | 5 |
| 2019 | FAST: enabling fast software/hardware prototype for network experimentationabstractThe evolution of new technologies in network community is getting ever faster. Yet it remains the case that prototyping those novel mechanisms on a real-world system (i.e. CPU-FPGA platforms) is both time and labor consuming, which has a serious impact on the research timeliness. In order to bring researchers out of trivial process in prototype development, this paper proposed FAST, a software hardware co-design framework for fast network prototyping. With the programming abstraction of FAST, researchers are able to prototype (using C, verilog or both) a wide spectrum of network boxes rapidly based on all kinds of CPU-FPGA platforms. FAST framework takes care of managing DMA, PCIe and Linux Kernel while providing a unified API for researchers so they can focus only on the packet processing functions. We demonstrate FAST framework's easy to use features with a number of prototypes and show we can get over 10x gains in performance or 1000x better accuracy in clock synchronization compared with their software versions. Xiangrui Yang 0002, Zhigang Sun 0002, Junnan Li 0002, Jinli Yan, Tao Li 0008, Wei Quan 0004, Donglai Xu, Gianni Antichi |
IWQoS | 2 |
| 2018 | Demonstration of Path-Based Packet Batcher for Accelerating Vectorized Packet ProcessingabstractRecently, a major challenge on generic multi-core network processing platforms is how to improve packet processing performance. Vector packet processor (VPP) is a modularized and high- performance software framework for building network dataplane applications. The key idea of VPP is to reduce instruction cache (i-cache) misses with vectorized packet processing. However, the packets in a vector may traverse different processing paths in some scenarios. In such case, the vector is split into several smaller vectors, and the per- packet overhead would increase. In this paper, we propose a Path-based Packet Batcher (PPB) to accelerate VPP. PPB is transparent to VPP, and it requires no modification to VPP. Before VPP processes packets, PPB batches the packets based on the processing paths they will traverse. We build a prototype based on FPGA to evaluate the performance optimizations to VPP with PPB. Experiment results show that the reduction of i-cache misses can be up to 57.6% when the batch size is 128. Jinli Yan, Tao Li 0008, Gaofeng Lv, Zhigang Sun 0002 |
SECON | 5 |
| 2018 | FAS: Using FPGA to Accelerate and Secure SDN Software SwitchesabstractSoftware-Defined Networking (SDN) promises the vision of more flexible and manageable networks but requires certain level of programmability in the data plane to accommodate different forwarding abstractions. SDN software switches running on commodity multicore platforms are programmable and are with low deployment cost. However, the performance of SDN software switches is not satisfactory due to the complex forwarding operations on packets. Moreover, this may hinder the performance of real-time security on software switch. In this paper, we analyze the forwarding procedure and identify the performance bottleneck of SDN software switches. An FPGA-based mechanism for accelerating and securing SDN switches, named FAS (FPGA-Accelerated SDN software switch), is proposed to take advantage of the reconfigurability and high-performance advantages of FPGA. FAS improves the performance as well as the capacity against malicious traffic attacks of SDN software switches by offloading some functional modules. We validate FAS on an FPGA-based network processing platform. Experiment results demonstrate that the forwarding rate of FAS can be 44% higher than the original SDN software switch. In addition, FAS provides new opportunity to enhance the security of SDN software switches by allowing the deployment of bump-in-the-wire security modules (such as packet detectors and filters) in FPGA. Wenwen Fu, Tao Li 0008, Zhigang Sun 0002 |
Secur. Commun. Networks | 3 |
| 2018 | OverWatch: A Cross-Plane DDoS Attack Defense Framework with Collaborative Intelligence in SDNabstractDistributed Denial of Service (DDoS) attacks are one of the biggest concerns for security professionals. Traditional middle-box based DDoS attack defense is lack of network-wide monitoring flexibility. With the development of software-defined networking (SDN), it becomes prevalent to exploit centralized controllers to defend against DDoS attacks. However, current solutions suffer with serious southbound communication overhead and detection delay. In this paper, we propose a cross-plane DDoS attack defense framework in SDN, called OverWatch, which exploits collaborative intelligence between data plane and control plane with high defense efficiency. Attack detection and reaction are two key procedures of the proposed framework. We develop a collaborative DDoS attack detection mechanism, which consists of a coarse-grained flow monitoring algorithm on the data plane and a fine-grained machine learning based attack classification algorithm on the control plane. We propose a novel defense strategy offloading mechanism to dynamically deploy defense applications across the controller and switches, by which rapid attack reaction and accurate botnet location can be achieved. We conduct extensive experiments on a real-world SDN network. Experimental results validate the efficiency of our proposed OverWatch framework with high detection accuracy and real-time DDoS attack reaction, as well as reduced communication overhead on SDN southbound interface. Biao Han 0003, Xiangrui Yang 0002, Zhigang Sun 0002, Jinshu Su |
Secur. Commun. Networks | 3 |
| 2018 | CSR: Classified Source Routing in Distributed NetworksabstractIn recent years cloud computing provides a new way to address the constraints of limited energy, capabilities, and resources. Distributed hash table (DHT) based distributed networks have become increasingly important for efficient communication in large-scale cloud systems. Previous studies mainly focus on improving the performance such as latency, scalability and robustness, but seldom consider the security demands on the routing paths, for example, bypassing untrusted intermediate nodes. Inspired by Internet source routing, in which the source nodes specify the routing paths taken by their packets, this paper presents CSR, a tag-based, Classified Source Routing scheme in distributed networks to satisfy the security demands on the routing paths. Different from Internet source routing which requires some map of the overall network, CSR operates in a distributed manner where nodes with certain security level are tagged with a label and routing messages requiring that level of security are forwarded only to the qualified next-hops. We show how this can be achieved efficiently, by simple extensions of the traditional routing structures, and safely, so that the routing is uniformly convergent. The effectiveness of our proposals is demonstrated through theoretical analysis and extensive simulations. Yiming Zhang 0003, Dongsheng Li 0001, Zhigang Sun 0002, Feng Zhao 0012, Jinshu Su, Xicheng Lu |
IEEE Trans. Cloud Comput. | 3 |
| 2017 | P5: Programmable Parsers with Packet-level Parallel Processing for FPGA-based SwitchesabstractThis paper presents P5, a programmable packet parser with packet-level parallel processing for FPGA-based switches. P5 overcomes both limitations. First, P5 has the programmability of dynamically updating parsing algorithms at run-time. Second, P5 exploits packet-level parallelism in the bottleneck of parsing pipeline to compensate FPGA's low clock frequency, and reduces resource consumption through a one-block recirculated strategy. Junnan Li 0002, Zhigang Sun 0002, Biao Han 0003 |
ANCS | 2 |
| 2017 | SDN-Based DDoS Attack Detection with Cross-Plane Collaboration and Lightweight Flow MonitoringabstractDistributed Denial of Service (DDoS) attacks are one of the biggest concerns for security professionals. Traditional DDoS attack detection mechanisms are based on middle-box devices or SDN controllers, which either lack network-wide monitoring information or suffer with serious southbound communication overhead and detection delay. In this paper, we propose a SDN-based DDoS attack detection framework with cross-plane collaboration called OverWatch, which performs a two-stage granularity filtering procedure between coarse-grained detection data plane and fine- grained detection control plane for abnormal flows. It leverages computational capabilities that currently underutilized on OpenFlow switches to shrink the detection range for fine-grained DDoS attack detections. In OverWatch, we propose a lightweight flow monitoring algorithm to capture the key features of DDoS attack traffics on the data plane by polling the values of counters in OpenFlow switches. Experiments are conducted in an evaluating network with a FPGA-based OpenFlow switch prototype and the Ryu controller, which reveal that our proposed OverWatch framework and flow monitoring algorithm can greatly improve the detection efficiency, as well as reduce the detection delay and southbound communication overhead. Xiangrui Yang 0002, Biao Han 0003, Zhigang Sun 0002 |
GLOBECOM | 3 |
| 2016 | Self-described buffer: A novel mechanism to improve packet I/O efficiency in LinuxabstractSocket buffer (SKB) is the standard data structure for exchanging packets and their control information between NIC driver and protocol stack. The overhead of dynamic SKB management has been considered as the significant bottleneck in packet I/O. Some novel non-SKB mechanisms, such as DPDK, were thus proposed to solve the problem. However, these mechanisms usually cannot be widely adopted in the data path of most packet forwarding applications, due to their incompatibility with SKB. In this paper, a new SKB-compatible mechanism, namely Self-described buffer (SDB), is proposed to improve the efficiency of packet I/O. SDB eliminates SKB allocation/deallocation overhead by offloading SKB management into NIC hardware. It also reduces the overhead of dynamic binding/unbinding operations existed in SKB management by statically binding related information in advance using the free space of Databuf. To evaluate the proposed approach, a SDB-enabled NIC and its driver has been designed and implemented based on FPGA. Experimental results show that the proposed SDB achieves 2× throughput compared with a traditional SKB mechanism in raw packet forwarding, and 34.75% improvement for typical network forwarding applications (e.g. IP forwarding, Bridge forwarding and SDN forwarding) on average. Jinli Yan, Zhigang Sun 0002, Tao Li 0008, Donglai Xu |
IWQoS | 3 |
| 2016 | BufferBank storage: an economic, scalable and universally usable in-network storage model for streaming data applications
Zhigang Sun 0002, Fei Yi, Jinshu Su |
Sci. China Inf. Sci. | 2 |
| 2016 | Efficient mismatched packet buffer management with packet order-preserving for OpenFlow networks
Jianbiao Mao, Biao Han 0003, Zhigang Sun 0002, Xicheng Lu |
Comput. Networks | 3 |
| 2016 | Design and implementation of Software Defined Hardware Counters for SDN
Tao Li 0008, Biao Han 0003, Zhigang Sun 0002 |
Comput. Networks | 4 |
| 2015 | FRINGE: Improving the scalability of Ethernet DCN via efficient software-defined edge controlabstractThis paper introduces a topology-independent software-defined edge control framework named FRINGE to scale out the Ethernet Datacenter Network (DCN). FRINGE exploits programmable OpenFlow-enabled switches deployed at the edge of DCN to aggregate the forwarding rules without introducing extra packet headers. We implement the proposed FRINGE framework in an SDN prototyping environment and validate it under three typical DCN topologies including Multi-Root Tree, HyperX and Jellyfish, where three different types of DCN workloads are applied. Evaluation results reveal that FRINGE can significantly reduce the total number of rules in all network devices and suppress most of the useless broadcast packets in the DCN. Jianbiao Mao, Biao Han 0003, Gaofeng Lv, Zhigang Sun 0002, Xicheng Lu |
IWQoS | 4 |
| 2015 | Towards high-performance packet processing on commodity multi-cores: current issues and future directions
Jinli Yan, Zhigang Sun 0002, Tao Li 0008, Minxuan Zhang |
Sci. China Inf. Sci. | 3 |
| 2014 | Demostration of Self-Described Buffer for Accelerating Packet Forwarding on Multi-core ServersabstractNetwork processing platform based on the multi-core CPU becomes more and more prevailing in nowadays. Buffer allocation/deallocation operations consume a large number of CPU cycles in packet I/O process. The problem becomes even worse in the scenario of packet forwarding, as buffer allocation/deallocation operations are more frequent than the host-based network applications. We thus propose a novel data structure for packet buffer management on multi-cores, named Self-Described Buffer (SDB), which merges the separated descriptor and metadata into packet buffer. SDB management overhead can be greatly reduced by utilizing the compact data structure, and zero-overhead buffer management can be further achieved by offloading SDB allocation/deallocation operations to NIC. We have prototyped SDB enabled NIC, named BcNIC, on NetFPGA-10G. In the demo, we will illustrate the advantages of the SDB scheme by comparing the performance of BcNIC with the traditional NIC on multi-core platforms. Zhigang Sun 0002, Tao Li 0008, Biao Han 0003, Gaofeng Lv |
CloudCom | 2 |
| 2014 | The Demonstration of Hyper Software Defined Hardware CountersabstractSoftware Defined Networking (SDN) provides efficient network and traffic management for data center network. As underlying devices in SDN, SDN switches must maintain a large number of hardware counters. Implementation of these counters faces serious challenges for SDN switches, i.e., High memory consumption and inflexibility. Thus, we previously proposed Software Defined Hardware Counters (SDHC), which decouples definition and implementation of counters to overcome these challenges. However, like traditional hardware counters, SDHC only supports passive statistical mode (i.e., The values of the counters can be only read passively by the controller). Based on the passive mode, most of applications need to send request messages at some frequency to obtain statistics, which causes some critical problems for SDN: i) low statistical accuracy, ii) high network bandwidth consumption. Hyper Software Defined Hardware Counters (Hyper SDHC) is thus proposed by extending SDHC. Through introducing the timer-triggering and updating-triggering statistics-reporting mechanisms, Hyper SDHC can naturally support active statistical mode, i.e., Counters actively report their values according to triggering condition. It can greatly enhance the statistical accuracy and reduce network bandwidth consumption between controller and switch. The demo of Hyper SDHC is implemented based on Net Magic platform. The demo will exhibit how Hyper SDHC works and how it supports a typical video quality monitor application. Tao Li 0008, Biao Han 0003, Zhigang Sun 0002 |
CloudCom | 4 |
| 2014 | Design of Software Defined hardware counters for SDNabstractImplementation of counters is a critical challenge for switches in today's Software-Defined Networking (SDN). In this paper, we address the current challenges in implementing SDN counters: high memory consumption, low utilization, and inflexibility. We introduce the concept of software defined hardware counters (SDHCs) for SDN. Our main idea is to make the switch-local CPU flexibly allocate memory space to each counter required by controllers. The ASIC of SDN switches transmits event records to the CPU, which contain updating information of the counters. Furthermore, the ASIC provides non-semantic counter memory space to be allocated by the CPU. Based on the proposed SDHC, an SDN controller can flexibly apply/release various counters for each counter category (e.g., each flow entry, each port) through the south-bound interface. It is shown that SDHC achieves high flexibility while reducing the memory space on ASIC. It also improves the update performance through alleviating the CPU overhead. Finally, we evaluate the performance of SDHC through comprehensive simulation study. Tao Li 0008, Biao Han 0003, Zhigang Sun 0002 |
LANMAN | 4 |
| 2014 | BufferBank: A distributed cache infrastructure for peer-to-peer application
Zhigang Sun 0002, Jianbiao Mao |
Peer-to-Peer Netw. Appl. | 2 |
| 2011 | Using NetMagic to observe fine-grained per-flow latency measurementsabstractWe introduce NetMagic to demonstrate the efficacy of RLI architecture RLI for the fine-grained per-flow latency measurements. In this demo, the main function of RLI is implemented in NetMagic, which is the key component of our experimental network comprising several computers and switches. We are going to show how NetMagic can provide rapid implementation and evaluation of RLI architecture that is difficult with commercial switch or router platforms. In the demo, the estimated fine-grained per-flow latency by RLI is monitored and dynamically presented. Further, the true latency with a resolution of 8ns is also provided by NetMagic for the evaluation. The efficacy of RLI architecture can be observed in a real-time fashion by the difference between estimated latencies and true ones. Tao Li 0008, Zhigang Sun 0002, Chunbo Jia, Myungjin Lee |
SIGCOMM | 2 |
| 2009 | E2EDSM: An Edge-to-Edge Data Service Model for Mass Streaming Media Transmission
Junfeng He, Ningwu He, Zhigang Sun 0002, Zhenghu Gong |
APPT | 4 |
| 2009 | Fast enumeration of maximal valid subgraphs for custom-instruction identificationabstractExtensible processors are increasingly becoming popular as they allow for incorporating custom instructions to meet design constraints. However, identifying custom instructions under architectural input/output ports constraint is a time consuming process particularly when large applications are considered. To rapidly identify the most profitable custom instructions with large inputs and outputs, this paper proposes a novel identification algorithm for enumerating maximal convex subgraphs containing no invalid node (i.e., maximal valid subgraphs). The proposed enumerating strategy is based on divide-and-conquer with a top-down manner, rather than the bottom-up manner utilized in the state-of-the-art. The division operation only considers invalid inner nodes of the given DFG, rather than taking all the invalid nodes into account, and thus accelerates enumeration of the maximal valid subgraphs. Experimental results show that, the improvement over the latest work is more than 90% for 60% DFG instances of the acknowledged benchmarks. Tao Li 0008, Zhigang Sun 0002, Jigang Wu, Xicheng Lu |
CASES | 2 |
| 2009 | MORT: A Technique to Improve Routing Efficiency in Fault-Tolerant Multipath RoutingabstractMultipath routing is thought of as a promising direction of the current routing system as it can improve the network performance in terms of reliability and throughput. However, there are some challenging problems to solve towards Internet-wide multipath routing. One of them is the dramatically increasing control message overhead caused by network dynamics. More message overhead will consume more computing resources and more storage. Meanwhile, more message overhead will lead to slower convergence process for routing protocols due to longer processing time. In this paper, we present MORT to solve the above problem. MORT is based on a technique called ¿information hiding¿. The ¿information hiding¿ technique allows routers in network to hide some routing information such as link failures and link cost changes to other routers without introducing any serious bad effect to the routing protocols. Multipath routing protocols embedded with MORT will have fewer routing message overhead and shorter routing convergence time when facing network events such as link failures and link recoveries. In the simulations, we apply MORT to a newly presented multipath protocol to show that MORT can reduce message overhead significantly as well as shortening the routing convergence time. Bin Dai 0001, Huabiao Lu, Zhigang Sun 0002, Ziming Song, Yanpeng Ma, Jinshu Su |
MSN | 3 |
| 2007 | A Uniform Fine-Grain Frame Spreading Algorithm for Avoiding Packet Reordering in Load-Balanced SwitchesabstractNetwork operators need high capacity router architectures that can offer scalability, provide throughput guarantees, and maintain packet ordering. However, current centralized crossbar-based architectures cannot scale to fast line rates and high port counts. On the other hand, while load-balanced switch architectures that rely on two identical stages of fixed configuration meshes appear to be an effective way to scale Internet routers to very high capacities, they incur a large worst-case packet reordering that is at best quadratic to the switch size. In this paper, we propose a Uniform Fine-grain Frame Spreading (UFFS) algorithm to avoid packet reordering throughout the load-balanced switch by assigning cells of the same flow to the fixed successive intermediate inputs. In order to distribute traffic equally among the intermediate inputs, a rotation mapping algorithm is used to construct a fixed mapping relationship between flows of different inputs and intermediate inputs in a round-robin fashion. The UFFS algorithm is distributed and can operate independently in each input. It spreads each flow to intermediate inputs according to the mapping relationship that is precomputed by the rotation mapping algorithm. We show that the UFFS algorithm can enforce packet ordering and achieve 100 % throughput with no additional communication of information among linecards. Jinshu Su, Zhigang Sun 0002, Jianbo Guan |
APSCC | 3 |