EDBT 2026 Demo / reviewers in the wild / expert
Zehua Guo 0001
dblp:145/8541
· DBLP profile ↗
103ranked-venue papers
19as first author
67since 2021 · last 2026
0000-0001-7314-410XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 79 · 16 first-author · 54 since 2021Systems, architecture and hardware · 15 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Artificial intelligence and machine learning · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Load-aware Ground Station Assignment for Low Earth Orbit Satellite NetworksabstractLow Earth Orbit (LEO) satellite constellations face ground segment bottlenecks due to uneven user demand, which overloads Ground-to-Satellite Links (GSLs). The common practice of routing traffic to the nearest Ground Station (GS) to minimize latency often causes severe load imbalance. This paper proposes a load-aware assignment strategy that minimizes the maximum GSL utilization by routing traffic to non-nearest GSs via inter-satellite links. To maintain service quality, assignments are constrained by a latency threshold relative to the nearest-GS baseline. We formulate this as a mixed-integer linear program. Preliminary results using realistic Starlink constellation parameters show the proposed solution can reduce average maximum GSL utilization, mitigating ground segment congestion. Songshi Dou, Jinxian Wu, Zehua Guo 0001, Kwan Lawrence Yeung |
CCNC | 3 |
| 2026 | QoS-driven Network Soft Slicing through Token-based Hierarchical Scheduling
Xiaoyang Fu, Zehua Guo 0001, Songshi Dou |
IWQoS | 2 |
| 2026 | Exploring Energy Saving Opportunities for Data Center Cooling Systems
Jianuo Li, Zehua Guo 0001, Xiaoyang Fu |
IWQoS | 2 |
| 2026 | Ensuring QoS Stability in Dynamic WANs through Selectively Optimized Range Routing
Zehua Guo 0001, Songshi Dou, Minghao Ye |
IWQoS | 2 |
| 2026 | CEINR: Critical Element-aware In-Network Recovery for Distributed Machine Learning
Jixing Yang, Shuran Zhang, Zehua Guo 0001, Songshi Dou, Jianye Wang |
IWQoS | 3 |
| 2026 | Fair Scheduling for Lossless RDMA Networks
Yazhu Zhao, Zehua Guo 0001, Xiaoyang Fu |
IWQoS | 2 |
| 2026 | Two-Level-Attention-Based Continuous Trajectory Design and Computation Offloading for Multi-UAV Cooperative Target Search
Haowen Zhu, Junpeng Hui, Zehua Guo 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2026 | Quest: Quality of Service-Centric Resilient Routing for Software-Defined Wide Area NetworksabstractEmerging network applications pose diverse Quality of Service (QoS) demands, prompting the adoption of Software-Defined Networking (SDN) in Wide Area Networks (WANs), known as Software-Defined Wide Area Networks (SD-WANs). SDN introduces path programmability, allowing controllers to dynamically adjust flow forwarding paths at their switches to accommodate varying traffic conditions and differentiated QoS requirements. However, controller failures can lead to the offline state of controlled switches, causing flows traversing these switches to lose path programmability and degrade QoS. Existing solutions typically focus on indirect metrics such as path programmability or coarse-grained load balancing, which often neglect the individual QoS demands of each flow and the control plane performance. To fill this research gap, we propose Quest, a QoS-aware resilient routing framework designed to preserve QoS in both data and control planes during controller failures. Compared to existing solutions, Quest explicitly addresses the critical limitations by jointly optimizing fine-grained QoS-aware flow routing and control plane latency. We formulate an optimization problem that jointly ensures QoS-aware flow routing and minimizes control plane delays. Due to the complexity of this problem, we utilize a linearization approach and further develop a heuristic QoS-centric resilient routing algorithm. Extensive simulations using real-world topologies and traffic traces demon-strate that Quest substantially improves network performance, achieving a 1.71% increase in average throughput ratio, a 46.65% reduction in average latency, and a 62.74% decrease in average control latency compared to baseline approaches. Songshi Dou, Zehua Guo 0001 |
IEEE Trans. Netw. | 2 |
| 2026 | Learning-Based Adaptive Range Routing for Traffic Engineering With Graph Neural NetworksabstractTraffic Engineering (TE) has been widely used by network operators to improve network performance and deliver better service quality. One major challenge for TE is providing routing strategies that can adapt to highly dynamic future traffic scenarios. Unfortunately, existing works either suffer severe performance degradation under unexpected traffic fluctuations, or sacrifice optimality to guarantee worst-case performance when traffic remains relatively stable. In this paper, we propose LARRI, a learning-based TE framework that predicts adaptive routing strategies for unknown future traffic scenarios. By integrating future demand range prediction and optimal range routing imitation into a single step, LARRI learns to generate a routing strategy that accommodates a wide range of possible future traffic matrices, thereby achieving a good trade-off between performance optimality and worst-case guarantees. Moreover, LARRI employs a scalable graph neural network architecture, which greatly facilitates both training and inference. Extensive simulations on six real-world network topologies show that LARRI achieves near-optimal load balancing in future traffic scenarios, improves worst-case performance by up to 43.3% over state-of-the-art baselines, and consistently provides the lowest end-to-end delay under dynamic traffic fluctuations. Minghao Ye, Junjie Zhang 0001, Zehua Guo 0001, H. Jonathan Chao |
IEEE Trans. Netw. | 3 |
| 2025 | Improving Transmission Flexibilty and Confidentiality using a Cloud-based VPN
Renzheng Wang, Tengteng Zhu, Zehua Guo 0001 |
APNet | 3 |
| 2025 | Critical Flow Range Routing for Wide Area Networks
Xiaoyang Fu, Minghao Ye, Zehua Guo 0001 |
APNet | 4 |
| 2025 | Improving In-Network Aggregation Efficiency with Multiple Hashing
Jixing Yang, Zehua Guo 0001 |
APNet | 2 |
| 2025 | A loss-based weighted aggregation method for federated learning with heterogeneous computing resources
Fan Yang 0020, Yi Yang 0068, Zehua Guo 0001 |
APNet | 5 |
| 2025 | Analyzing the Delay Bound for Network SlicingabstractQueuing delay is a critical part of end-to-end delay. Existing delay analysis methods in network devices either oversimplify switch scheduling models or are dependent on traffic statistics. To precisely analyze the queue delay, we propose a new analysis method, which has two features: (1) incorporating switch-specific scheduling architectures and algorithms and (2) combining deterministic and stochastic parameter analysis. We take network slicing as an example to test our method. Experiments show that our proposed method reduces the average delay estimation error by over 80% compared to the traditional delay analysis method for network slicing and captures the actual trend. Xiaoyang Fu, Yazhu Zhao, Rongfei Zeng, Zehua Guo 0001 |
IWQoS | 5 |
| 2025 | Maintaining Predictable Traffic Engineering Performance Under Controller Failures for Software-Defined WANsabstractMany new cloud services and applications have emerged recently. They account for a large share of traffic in Wide Area Networks (WANs) and provide traffic with various Quality of Service (QoS) requirements. Software-Defined Wide Area Network (SD-WAN) offers a promising opportunity for improving the performance of these applications with flexible network management. Nevertheless, SD-WANs are managed by controllers, and unpredictable controller failures may degrade flexible network management. Switches previously controlled by the failed controllers become offline, and flows traversing these offline switches lose the path programmability to route flows on available forwarding paths. Thus, these offline flows cannot be routed/rerouted on available paths to accommodate potential traffic variations, leading to severe performance degradation. Traffic Engineering (TE) is a prevalent network application, which aims to enable differentiable QoS for these numerous cloud services and applications. However, TE performance cannot be guaranteed when controller failures happen due to the loss of flexible network management. Existing recovery solutions reassign offline switches to other active controllers to recover the degraded path programmability but may not promise good TE performance since higher path programmability does not necessarily guarantee satisfactory TE performance. In this paper, we propose ARES to provide predictable TE performance under controller failures. We formulate an optimization problem, which aims to maintain predictable TE performance by jointly considering fine-grained flow-controller reassignment and flow rerouting. Given that the proposed problem is proven to be NP-hard, we further propose a heuristic algorithm to efficiently solve this problem. Specifically, when controller failures occur, ARES updates real-time network information with traffic traces and failure status to calculate optimal flow-controller reassignment and flow rerouting policies. ARES then reassigns and reroutes offline flows to maintain predictable TE performance. Extensive simulation results under two real-world topologies with traffic traces demonstrate that our problem formulation exhibits comparable load balancing performance to optimal TE solution without controller failures, and the proposed ARES can significantly improve average load balancing performance by up to 35.79% with low computation time compared with the state-of-the-art solution. Songshi Dou, Zehua Guo 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2025 | Path-Based Graph Neural Network for Robust and Resilient Routing in Distributed Traffic EngineeringabstractDistributed Traffic Engineering (TE) aims to optimize network performance by generating individual routing strategies at each router without a global view of the network. A major challenge for these TE solutions is handling performance degradation caused by unexpected traffic fluctuations and unpredictable link failures. Recently, Machine Learning (ML) techniques have introduced new opportunities to enhance distributed TE. In this paper, we propose Path-Based Graph Neural Network (PathGNN), which leverages the emerging GNN architecture to quickly infer robust and resilient routing strategies in a distributed manner to accommodate unexpected network conditions. PathGNN adopts a novel path-link bipartite graph modeling approach to capture the dynamics of link resources shared by routing paths. It then performs efficient GNN message exchanges among routers to make adaptive local routing decisions for better load balancing. Additionally, PathGNN leverages Supervised Learning (SL) to directly learn from optimal routing strategies through efficient offline training. Evaluation results on four real-world network topologies demonstrate PathGNN’s strong generalization capability. Compared to state-of-the-art distributed TE solutions, PathGNN improves the load balancing performance by at least 24.4% with lower end-to-end delay under dynamic traffic scenarios, and also boosts performance by up to 35.3% under multiple link failures. Minghao Ye, Junjie Zhang 0001, Zehua Guo 0001, H. Jonathan Chao |
IEEE J. Sel. Areas Commun. | 3 |
| 2025 | Unleashing the Potential of LEO Constellations in Building Resilient and Low-Latency Control Plane for SD-WANsabstractDelivering seamless network services (e.g., video streaming, AR/VR, and cloud gaming) in Wide Area Networks (WANs) relies on flexible traffic management, which is facilitated by Software-Defined Wide Area Networks (SD-WANs), or SDN in WANs. In SD-WANs, control and data traffic usually share the same links for cost savings, a practice known as in-band control. The SDN controller can periodically send control messages to switches via control channels for routing policy updates to maintain satisfactory network performance. However, when a link failure occurs, control channels become disconnected. Since the controller can no longer communicate with the switches, the desired flexible traffic management cannot be promised. Although backup paths can be preconfigured to reconnect some control channels, high control latency may be introduced due to the meandering routes. Fortunately, commercial Low-Earth Orbit (LEO) mega-constellations, which provide pervasive and low-latency Internet services, present a promising solution to this control resiliency issue. Inspired by the rapid deployment of these LEO constellations, we propose a novel control plane design calledSpaceHelperto leverage the LEO satellite network for improving SD-WANs’ control resiliency.SpaceHelpersmartly integrates the LEO satellite network with terrestrial SD-WAN to reconnect control channels during link failure, which is formulated as an optimization problem to minimize overall control latency. A heuristic algorithm is proposed to solve the problem efficiently and guarantee prompt control channel reconnection. Performance evaluations are conducted using the Starlink constellation and real-world WAN topologies. Compared to the state-of-the-art in-band solution, we show thatSpaceHelpercan not only provide 100% resiliency, but also significantly reduce average control latency by up to 72.6% and 70.2% under GÉANT and Abilene topologies, respectively. Songshi Dou, Zehua Guo 0001, Kwan Lawrence Yeung |
IEEE Trans. Netw. | 2 |
| 2025 | Revisiting the In-Network Aggregation in Distributed Machine LearningabstractDistributed Machine Learning (DML) is proposed to accelerate machine learning model training by utilizing multiple training nodes to train models in parallel. Recent studies apply emerging In-Network Aggregation (INA) to further improve training efficiency by offloading the gradient aggregation process from hosts to programmable switches. However, existing INA solutions neither provide high training performance due to inefficient gradient aggregation with a single switch nor are easily deployed because of modifying the protocol stack in hosts. In this paper, we propose an easily deployable INA-based solution called Hierarchical In-Network Aggregation (HINA) to accelerate DML training process by hierarchically performing multiple aggregations in the Data Center Network (DCN). We formulate the gradient aggregation problem as the Joint Gradient Routing and Sending Rate (JGRSR) problem, which is a Mixed Integer Linear Programming (MILP) problem with high computation complexity. In addition, we propose HINA using progressive rounding and randomized rounding to determine the paths of gradient flows and the sending rates of training nodes to simplify and solve the JGRSR problem. Simulation results show that HINA reduces communication time by 61%-92% and decreases network load by 39%-66%, compared with state-of-the-art solutions. Haowen Zhu, Zehua Guo 0001 |
IEEE Trans. Netw. | 2 |
| 2025 | DINA: Toward Determined In-Network Aggregation for Distributed Machine LearningabstractDistributed Machine Learning (DML) utilizes parallel computation on multiple training nodes to accelerate machine learning model training. Parameter Server (PS) is a typical DML enabler and is widely used in industry and academia. Existing works propose to apply the emerging In-Network Aggregation (INA) technique to improve model training efficiency by offloading the whole gradient aggregation process in PS from hosts to programmable switches. However, existing INA systems may suffer from undetermined model training efficiency and service quality, given that many gradient aggregation processes are still performed by the server under irrational gradient aggregation strategies. In this paper, we propose a Deterministic In-Network Aggregation (DINA) scheme to improve model training efficiency by enhancing the efficiency of INA utilization in DML. Our key observation is to further increase worker sending rates by reducing gradient packets’ RTT (i.e., realizing packet sub-RTT). Based on this observation, DINA rationally selects the optimal global gradient aggregation switch depending on the switches’ available memory, worker sending rate, and server processing capacity. As a result, DINA reduces the dependence of INA systems on the server, improves worker sending rates, and mitigates network traffic load. We formulate the sub-RTT-INA-based gradient aggregation problem as a mixed-integer nonlinear programming problem. To efficiently solve the problem, we simplify it by transforming the nonlinear constraints into linear constraints and propose a mixed solution that combines randomized rounding and heuristic mechanisms. Simulation results show that DINA can provide determined training by reducing communication time by 12%-17% and network load by 28%-50% compared with existing solutions, thus taking full advantage of INA and realizing a determined INA service. Haowen Zhu, Zehua Guo 0001, Minghao Ye |
IEEE Trans. Netw. | 2 |
| 2024 | Exploring the Impact of Traffic Scheduling on Network Soft SlicingabstractNetwork slicing is a promising technique to enable a physical network to support various applications with different demands for network services. In this paper, we propose a traffic scheduler for network soft slicing and evaluate it in a typical soft-slicing case. Xiaoyang Fu, Zehua Guo 0001 |
APNet | 2 |
| 2024 | Halflife: An Adaptive Flowlet-based Load Balancer with Fading Timeout in Data Center NetworksabstractModern data centers (DCs) employ various traffic load balancers to achieve high bisection bandwidth. Among them, flowlet switching has shown remarkable performance in both load balancing and upper-layer protocol (e.g., TCP) friendliness. However, flowlet-based load balancers suffer from the inflexibility of flowlet timeout value (FTV) and result in sub-optimal performance under various application workloads. To this end, we propose Halflife, a novel flowlet-based load balancer that leverages fading FTVs to reroute traffic promptly under different workloads without any prior knowledge. Halflife not only balances traffic better, but also avoids the performance degradation caused by frequent oscillation or shifting of lows between paths. Furthermore, Halflife's fading mechanism is not only compatible with most flowlet-based load balancers, such as CONGA and LetFlow, but also improves their performance when leveraging flowlet switching in RDMA network. Through testbed experiments and simulations, we prove that Halflife improves the performance of CONGA and LetFlow by 10% ~ 150%, and it outperforms other load balancers by 30% ~ 200% across most application workloads. Sen Liu 0002, Yongbo Gao, Jiarui Ye, Furong Liang, Zerui Tian, Quanwei Sun, Zehua Guo 0001, Yang Xu 0010 |
EuroSys | 10 |
| 2024 | An ML-Accelerated Framework for Large-Scale Constrained Traffic EngineeringabstractTraffic engineering (TE) mechanisms are crucial for achieving optimal levels of network performance over wide-area networks across geographically distributed datacenters. Existing work on traffic engineering formulated the challenges at hand as combinatorial optimization problems, which could take hours to compute on modern wide-area network topologies at the scale of thousands of nodes. To improve the performance of TE mechanisms, we introduce DeepTE, a new TE framework based on machine learning (ML) that is designed for the best possible scalability and performance, capable of completing the computation within milliseconds with networks involving thousands of nodes, and of generating near-optimal TE policies while guaranteeing that all constraints are satisfied. DeepTE is also designed with a distributed ML model architecture, which can be horizontally scaled up to multiple GPUs for even better performance. With real-world traffic matrices, our extensive array of performance evaluations of DeepTE on various network topologies and TE problems show that DeepTE is capable of producing policies within 5% of the optimal results while offering up to 100x performance improvements over state-of-the-art traffic engineering mechanisms. Ben Hok Ng, Qiao Xiang, Zehua Guo 0001 |
ICDCS | 5 |
| 2024 | Hierarchical Sketch: An Efficient, Scalable and Latency-aware Content Caching Design for Content Delivery NetworksabstractContent Delivery Networks (CDNs) are designed to reduce user-perceived waiting times and alleviate backbone bandwidth pressure. Since CDN cache servers have limited storage capacity, effective cache replacement policies are needed. However, existing CDN cache replacement policies mainly focus on improving content hit rates. As a result, some content with long origin fetch latency may not be cached, resulting in the long tail latency and degrading user experience. In this paper, we present Hierarchical Sketch, an efficient, scalable, and latency-aware cache replacement algorithm. Our approach leverages hierarchical slicing and voting mechanisms on a modified sketch to optimize content caching, reducing sorting complexity from O(log n) to O(1) with minimal loss of hit rate. Extensive simulations on synthetic and real-life industry CDN traces demonstrate that Hierarchical Sketch outperforms other algorithms in four different scenarios, with up to a 15% improvement. Huifeng Xing, Yuyan Ding, Huiru Huang, Sen Liu 0002, Zehua Guo 0001, Muath Al-Hasan, Mohamed Adel Serhani, Yang Xu 0010 |
IWQoS | 6 |
| 2024 | Toward Determined Service for Distributed Machine LearningabstractParameter Server (PS) is a typical Distributed Machine Learning (DML) enabler and widely used in industry and academia. Existing works propose to apply the emerging In-Network Aggregation (INA) technique to improve model training efficiency. However, existing INA systems may suffer from undetermined model training efficiency and service quality, given that many gradient aggregation processes are still performed by the server under irrational gradient aggregation strategies. In this paper, we propose a Deterministic In-Network Aggregation (DINA) scheme to improve model training efficiency by enhancing the efficiency of INA utilization in DML. Our key observation is to further increase worker sending rates by reducing gradient packets’ RTT. Based on this observation, DINA can rationally select the optimal global gradient aggregation switch depending on the switches’ available memory, worker sending rate, and server processing capacity. Simulation results show that DINA can provide determined training by improving worker sending rates by 16%-87% and network load by 28%-46.8% compared with existing solutions. Haowen Zhu, Minghao Ye, Zehua Guo 0001 |
IWQoS | 3 |
| 2024 | ARES: Predictable Traffic Engineering under Controller Failures in SD-WANsabstractEmerging web applications (e.g., video streaming and Web of Things applications) account for a large share of traffic in Wide Area Networks (WANs) and provide traffic with various Quality of Service (QoS) requirements. Software-Defined Wide Area Networks (SD-WANs) offer a promising opportunity to enhance the performance of Traffic Engineering (TE), which aims to enable differentiable QoS for numerous web applications. Nevertheless, SD-WANs are managed by controllers, and unpredictable controller failures may undermine flexible network management. Switches previously controlled by the failed controllers may become offline, and flows traversing these offline switches lose the path programmability to route flows on available forwarding paths. Thus, these offline flows cannot be routed/rerouted on previous paths to accommodate potential traffic variations, leading to severe TE performance degradation. Existing recovery solutions reassign offline switches to other active controllers to recover the degraded path programmability but fail to promise good TE performance since higher path programmability does not necessarily guarantee satisfactory TE performance. In this paper, we propose ARES to provide predictable TE performance under controller failures. We formulate an optimization problem to maintain predictable TE performance by jointly considering fine-grained flow-controller reassignment using P4 Runtime and flow rerouting and propose ARES to efficiently solve this problem. Extensive simulation results demonstrate that our problem formulation exhibits comparable load balancing performance to optimal TE solution without controller failures, and the proposed ARES significantly improves average load balancing performance by up to 43.36% with low computation time compared with existing solutions. Songshi Dou, Zehua Guo 0001 |
WWW | 3 |
| 2024 | Mitigating the impact of controller failures on QoS robustness for software-defined wide area networks
Songshi Dou, Zehua Guo 0001 |
Comput. Networks | 3 |
| 2024 | Byzantine-robust Federated Learning via Cosine Similarity Aggregation
Tengteng Zhu, Zehua Guo 0001, Jiaxin Tan, Songshi Dou, Wenrun Wang, Zhenzhen Han |
Comput. Networks | 2 |
| 2024 | Maintaining the Network Performance of Software-Defined WANs With Efficient Critical RoutingabstractSoftware-Defined Networking (SDN) brings new opportunities to improve network performance of Wide Area Networks (WANs). To enhance the control plane’s processing ability, a Software-Defined Wide Area Network (SD-WAN) usually employs multiple SDN controllers to control a large scale network. The controllers keep a consistent network state with each other via the controller synchronization. A controller synchronization usually involves all controllers, and the network cannot operate until the synchronization is done. Therefore, existing controller synchronization schemes could affect network performance and increase high resource consumption, thus increasing the complexity of running the SD-WAN. In this paper, we propose ThresHold-based critical flOw Routing (THOR) to maintain good load balancing performance of the network. THOR identifies and reroutes critical flows in single domains to achieve the network requirement without affecting other domains. If the rerouting does not satisfy the performance threshold, we generate and reroute several traversing flows among several domains to improve network performance. Simulation results show that THOR achieves 87% of the load balancing performance and reduces the number of synchronizations by approximately 80%, compared with the existing optimal flow routing. Zehua Guo 0001, Songshi Dou, Bida Zhang, Weichao Wu |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2024 | Realizing the Carbon-Aware Service Provision in ICT SystemabstractThe ever-growing carbon emission of information infrastructure accounts for a significant proportion of the global carbon emissions. Existing studies reduce carbon consumption mainly by improving power efficiency on specific facilities or energy source structures. However, these methods do not jointly consider the impact of computation and network resource distribution on carbon emission. In this paper, we propose a data-driven scheme named EcoNet using reinforcement learning to reduce carbon emissions by jointly scheduling computation and network resources. We dynamically monitor the status of the computation and network facilities using cloud-edge collaboration and software-defined networking. Based on the collected status information, we formulate the resource scheduling problem as an optimization problem, which comprehensively considers the carbon emission, electricity price, and quality of service. The problem has high computation complexity, and we solve the problem with the proposed EcoNet to achieve efficient scheduling and near-optimal performance based on the collected network status information. The evaluation results show that EcoNet can maintain good Quality of Service and save at least 17% of the overall cost considering the electricity bills and carbon emissions. Penghao Sun, Julong Lan, Yuxiang Hu 0001, Zehua Guo 0001, Jiangxing Wu 0001 |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2024 | EPIC: Traffic Engineering-Centric Path Programmability Recovery Under Controller Failures in SD-WANsabstractSoftware-Defined Wide Area Networks (SD-WANs) offer a promising opportunity to enhance the performance of Traffic Engineering (TE). With the help of Software-Defined Networking (SDN), TE can promptly respond to traffic changes and maintain network performance by leveraging a global network view. One of the key benefits of SDN for TE is path programmability, which is empowered by SDN controllers to enable dynamic adjustments of flows’ forwarding paths. However, controller failures pose new challenges for SD-WANs since path programmability could be decreased due to the increasing number of offline flows, leading to potential TE performance degradation. Existing recovery solutions mainly focus on recovering path programmability for improving unpredictable network performance but cannot guarantee consistently satisfactory TE performance as expected, since path programmability can only indirectly evaluate network performance. In this paper, we propose EPIC to ensure robust TE performance under controller failures. We observe that frequently rerouted flows could greatly influence TE performance. Enlightened by this, EPIC introduces a novel metric called the TE performance-centric ratio to assess the relevance of different path programmability values for TE performance. The key idea of EPIC lies in identifying frequently rerouted flows during TE operations and prioritizing recovery of the path programmability of these flows under controller failures. We formulate an optimization problem to maximize TE performance-centric path programmability and propose an efficient heuristic algorithm to solve this problem. Evaluation results demonstrate that EPIC can improve average load balancing performance by up to 55.6% compared with baselines. Songshi Dou, Jianye Wang, Zehua Guo 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2024 | Toward Improved Path Programmability Recovery for Software-Defined WANs Under Multiple Controller FailuresabstractEnabling path programmability is an essential feature of Software-Defined Networking (SDN). During controller failures in Software-Defined Wide Area Networks (SD-WANs), a resilient design should maintain path programmability for offline flows, which were controlled by the failed controllers. Existing solutions can only partially recover the path programmability rooted in two problems: 1) the implicit preferable recovering flows with long paths and 2) the sub-optimal remapping strategy in the coarse-grained switch level. In this paper, we propose ProgrammabilityGuardian to recover the path programmability of offline flows while maintaining low communication overhead. These goals are achieved through the fine-grained flow-level mappings enabled by existing SDN techniques. ProgrammabilityGuardian configures the flow-controller mappings to recover offline flows with a similar path programmability, maximize the total programmability of the offline flows, and minimize the total communication overhead for controlling these recovered flows. Simulation results of different controller failure scenarios under two different topologies show that ProgrammabilityGuardian recovers offline flows with a balanced path programmability, improves the total programmability of the recovered flows up to 68% and 70%, and reduces the communication overhead by 96% and 99%, compared with the baseline algorithm. Zehua Guo 0001, Songshi Dou, Wenchao Jiang, Yuanqing Xia |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Maintaining Control Resiliency for Traffic Engineering in SD-WANsabstractSoftware-Defined Networking (SDN) is introduced to Wide Area Networks (WANs) to facilitate network operations and management. One main advantage of SDN is path programmability, i.e., the ability to change forwarding path of flows by controlling underlying SDN switches to accommodate traffic variations. However, the SDN controllers may experience unexpected failure and thus lose its path programmability. The typical solution is to let active controllers control offline switches, which are controlled by failed controllers, by establishing the remapping between the offline switches and active controllers. However, existing remapping solutions do not consider the impact of controller failure on network performance and cannot exhibit predictable network performance under controller failure. In this paper, we take Traffic Engineering (TE) as a typical network scenario and propose Traffic Engineering-Aware Controller-switcH remapping rEcoveRy named TEACHER. We introduce Traffic-aware Path Programmability (TPP) as a new metric to describe the impact of controller failure on TE and design TEACHER based on this metric to smartly recover offline switches. Simulation results show that TEACHER can increase the overall TPP by up to 83.0% and improves the load balancing performance by up to 58.2% under Sprintlink topology with relatively low computation time, compared with baselines. Zehua Guo 0001, Songshi Dou, Jiawei Weng, Xiaoyang Fu, Yuanqing Xia |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Prophet: Traffic Engineering-Centric Traffic Matrix PredictionabstractTraffic Matrix (TM), which records traffic volumes among network nodes, is important for network operation and management. Due to cost and operation issues, TMs cannot be directly measured and collected in real time. Therefore, many studies work on predicting future TMs based on historical TMs. However, existing works are usually accuracy-centric prediction solutions that mainly focus on improving predicting accuracy of flows’ sizes (i.e., values of elements in TMs) without considering the practical application of TMs. In this paper, we propose a novel TM prediction solution called Prophet for Traffic Engineering (TE), a typical application for TMs which takes TMs as input to optimize routing. We identify that the critical property (i.e., ratio among elements) in a TM plays an important role in TE’s performance. Based on this analysis, we adopt the matrix normalization to maintain the critical property in TMs and customize a TE-centric angle loss function to introduce scale invariance of TMs for capturing the overall relationship error. Different from the element-wise Mean Squared Error (MSE) loss function in accuracy-centric prediction solutions, our proposed TE-centric angle loss function has a clear geometric interpretation, which confines the angle between predicted TM and real TM to zero. Simulation results show that the predicted TMs from Prophet can improve the performance of link-level TE and path-level TE by up to 45.4% and 52.8%, respectively, compared to existing solutions. Yuntian Zhang, Tengteng Zhu, Junjie Zhang 0001, Minghao Ye, Songshi Dou, Zehua Guo 0001 |
IEEE/ACM Trans. Netw. | 7 |
| 2023 | Efficient and Structural Gradient Compression with Principal Component Analysis for Distributed TrainingabstractDistributed machine learning is a promising machine learning approach for academia and industry. It can generate a machine learning model for dispersed training data via iterative training in a distributed fashion. To speed up the training process of distributed machine learning, it is essential to reduce the communication load among training nodes. In this paper, we propose a layer-wise gradient compression scheme based on principal component analysis and error accumulation. The key of our solution is to consider the gradient characteristics and architecture of neural networks by taking advantage of the compression ability enabled by PCA and the feedback ability enabled by error accumulation. The preliminary results on image classification task show that our scheme achieves good performance and reduces 97% of the gradient transmission. Jiaxin Tan, Zehua Guo 0001 |
APNet | 3 |
| 2023 | Roracle: Enabling Lookahead Routing for Scalable Traffic Engineering with Supervised LearningabstractTraditional Traffic Engineering (TE) usually balances the load on network links by formulating and solving a routing optimization problem based on measured Traffic Matrices (TMs). Given that traffic demands could change unexpectedly and significantly in realistic scenarios, routing strategies opti-mized based on currently measured TMs might not work well in future traffic scenarios. To compensate for the mismatch between stale routing decisions and future TMs, network operators may perform routing updates more frequently, which could introduce significant network disturbance and service disruption. Moreover, given the high routing computation overhead of TE optimization in today's large-scale networks, routing updates could experience severe delay and thus cannot accommodate future traffic changes in time. To address these challenges, we propose Roracle, a scalable learning-based TE that quickly predicts a good routing strategy for a long sequence of future TMs, while the learning process is guided by the optimal solutions of Linear Programming (LP) problems using Supervised Learning (SL). We design a scalable Graph Neural Network (GNN) architecture that greatly facilitates training and inference processes to accelerate TE in large networks. Extensive simulation results on real-world network topologies and traffic traces show that Roracle outperforms existing TE solutions by up to 36% in terms of worst-case performance under future unknown traffic scenarios. Additionally, Roracle achieves good scalability by providing at least$71\times$speedup over the most efficient baseline method in large-scale networks. Minghao Ye, Junjie Zhang 0001, Zehua Guo 0001, H. Jonathan Chao |
ICNP | 3 |
| 2023 | LARRI: Learning-based Adaptive Range Routing for Highly Dynamic Traffic in WANsabstractTraffic Engineering (TE) has been widely used by network operators to improve network performance and provide better service quality to users. One major challenge for TE is how to generate good routing strategies adaptive to highly dynamic future traffic scenarios. Unfortunately, existing works could either experience severe performance degradation under unexpected traffic fluctuations or sacrifice performance optimality for guaranteeing the worst-case performance when traffic is relatively stable. In this paper, we propose LARRI, a learning-based TE to predict adaptive routing strategies for future unknown traffic scenarios. By learning and predicting a routing to handle an appropriate range of future possible traffic matrices, LARRI can effectively realize a trade-off between performance optimality and worst-case performance guarantee. This is done by integrating the prediction of future demand range and the imitation of optimal range routing into one step. Moreover, LARRI employs a scalable graph neural network architecture to greatly facilitate training and inference. Extensive simulation results on six real-world network topologies and traffic traces show that LARRI achieves near-optimal load balancing performance in future traffic scenarios with up to 43.3% worst-case performance improvement over state-of-the-art baselines, and also provides the lowest end-to-end delay under dynamic traffic fluctuations. Minghao Ye, Junjie Zhang 0001, Zehua Guo 0001, H. Jonathan Chao |
INFOCOM | 3 |
| 2023 | RateSheriff: Multipath Flow-aware and Resource Efficient Rate Limiter Placement for Data Center NetworksabstractEmerging cloud services and applications request different Quality of Service (QoS) in Data Center Networks (DCNs). To meet these various requirements, programmable switch-based rate limiters are introduced to provide performance isolation and benefit from easy control and fast deployment. However, existing programmable switch-based rate limiters have two limitations: (1) multipath flows (i.e., MultiPath TCP) cannot be precisely limited, and (2) rate limiter placement solutions in DCNs are missing. These limitations could lead to poor rate limiting performance and low bandwidth utilization. In this paper, we propose RateSheriff to improve rate limiting performance by providing multipath flow-aware and resource efficient rate limiter placement for programmable switch-enabled DCNs. We identify and associate subflows to a multipath flow by extracting and comparing specific packets and header fields. By solving the formulated resource efficient rate limiter placement problem, we can improve rate limiting performance and balance memory utilization among programmable switches in DCNs. Simulation results show that RateSheriff can correctly limit the rate of multipath flows, improve rate limiting performance by up to 46%, and improve memory balancing performance by up to 79% with low computation time, compared with baselines. Songshi Dou, Yongchao He, Sen Liu 0002, Wenfei Wu, Zehua Guo 0001 |
IWQoS | 5 |
| 2023 | Reinforcement Learning-based Traffic Engineering for QoS Provisioning and Load BalancingabstractEmerging applications pose different Quality of Service (QoS) requirements for the network, where Traffic Engineering (TE) plays an important role in QoS provisioning by carefully selecting routing paths and adjusting traffic split ratios on routing paths. To accommodate diverse QoS requirements of traffic flows under network dynamics, TE usually periodically computes an optimal routing strategy and updates a significant number of forwarding entries, which introduces considerable network operation management overhead. In this paper, we propose QoS-RL, a Reinforcement Learning (RL)-based TE solution for QoS provisioning and load balancing with low management overhead and service disruption during routing updates. Given the traffic matrices that represent the traffic demands of high and low priority flows, QoS-RL can intelligently select and update only a few destination-based forwarding entries to satisfy the QoS requirements of high priority traffic while maintaining good load balancing performance by rerouting a small portion of low priority traffic. Extensive simulation results on four real-world network topologies demonstrate that QoS-RL provides at least 95.5 % of optimal end-to-end delay performance on average for high priority flows, and also achieves above 90 % of optimal load balancing performance in most cases by updating only 10% of destination-based forwarding entries. Minghao Ye, Junjie Zhang 0001, Zehua Guo 0001, H. Jonathan Chao |
IWQoS | 4 |
| 2023 | SMAF: a secure and makespan-aware framework for executing serverless workflows
Zehua Guo 0001, Guozhen Chen |
Sci. China Inf. Sci. | 3 |
| 2023 | Congestion-Aware Critical Gradient Scheduling for Distributed Machine Learning in Data Center NetworksabstractDistributed Machine Learning (DML) is proposed not only to accelerate the training of machine learning, but also to solve the inadequate ability for handling a large amount of training data. It adopts multiple computing nodes in data center to collaboratively work in parallel at the cost of high communication overhead. Gradient Compression (GC) is introduced to reduce the communication overhead by reducing the number of synchronized gradients among computing nodes. However, existing GC solutions suffer from varying network congestion. To be specific, when some computing nodes experience high network congestion, their gradient transmission process could be significantly delayed, slowing down the entire training process. To solve the problem, we propose FLASH, a congestion-aware GC solution for DML. FLASH accelerates the training process by jointly considering the iterative approximation of machine learning and dynamic network congestion scenarios. It can maintain good training performance by adaptively adjust and schedule the number of synchronized gradients among computing nodes. We evaluate the effectiveness of FLASH using AlexNet and Resnet18 under different network congestion scenarios. Simulation results show that under the same number of training epochs, FLASH reduces training time 22-71%, maintains good accuracy, and low loss, compared with the existing memory top-K GC solution. Zehua Guo 0001, Sen Liu 0002, Jineng Ren, Yang Xu 0010 |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | SpongeTraining: Achieving High Efficiency and Accuracy for Wireless Edge-Assisted Online Distributed LearningabstractEdge-assisted Distributed Learning (EDL) is a popular machine learning paradigm that uses a set of distributed edge nodes to collaboratively train a machine learning model using training data. Most of existing works implicitly assume that the fixed amount of training data is pre-collected and dispatched from user devices to edge nodes. In real world, however, training data in edge nodes are collected from user devices through wireless networks, and the volume and distribution of training data in edge nodes could exhibit temporal and spatial fluctuations due to varying wireless situations (e.g., network congestion, link capacity variation). In this way, existing solutions suffer from slow convergence and low accuracy. In this paper, we propose SpongeTraining to achieve high efficiency and accuracy for online EDL. To accommodate to fluctuations in training data, SpongeTraining uses a buffer at each worker to store received training data and adaptively adjusts training batch size and learning rate of each worker based on training data extracted from the buffer. Experiment results based on real-world datasets show that SpongeTraining outperforms existing solutions by accelerating the training process up to 50% for reaching the same training accuracy. Zehua Guo 0001, Sen Liu 0002, Jineng Ren, Yang Xu 0010, Yi Wang 0004 |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Exploring the Impact of Critical Programmability on Controller Placement for Software-Defined Wide Area NetworksabstractControl latency is a critical concern for deploying Software-Defined Networking (SDN) into Wide Area Networks (WANs). A Software-Defined WAN (SD-WAN) can be divided into multiple domains controlled by multiple controllers with a logically centralized view. The control latency is related to the placement of controllers and mappings between switches and controllers. Existing solutions usually consider the propagation delay between switches and controllers as the evaluation metric and fail to consider many important factors of dynamic network states. In this paper, we propose ProgrammabilityExplorer (PE) to optimize the control latency in SD-WAN. Inspired by the selection of critical flows, which have a critical impact on network performance, PE considers the programmability of critical flows at switches and uses this metric to decide the placement of controllers and mappings between switches and controllers. Simulation results show that PE can reduce the control latency by up to 62.3%, 27.5%, 58.3%, and 61.7% under GÉANT, Abilene, Sprintlink, and Tiscali topologies respectively, compared with baseline algorithms. Songshi Dou, Zehua Guo 0001 |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | Toward Flexible and Predictable Path Programmability Recovery Under Multiple Controller Failures in Software-Defined WANsabstractSoftware-Defined Networking (SDN) promises good network performance in Wide Area Networks (WANs) with the logically centralized control using physically distributed controllers. In Software-Defined WANs (SD-WANs), maintaining path programmability, which enables flexible path change on flows, is crucial for maintaining network performance under traffic variation. However, when controllers fail, existing solutions are essentially coarse-grained switch-controller mapping solutions and only recover the path programmability of a limited number of offline flows, which traverse offline switches controlled by failed controllers. In this paper, we propose FlexibleProgrammabilityMedic (FlexPM) to provide predictable path programmability recovery under multiple controller failures in SD-WANs. The key idea of FlexPM is to approximately realize flow-controller mappings using hybrid SDN/legacy routing supported by high-end commercial SDN switches. Using the hybrid routing, we can recover programmability by selecting a routing mode for each offline flow at each offline switch in a fine-grained way to fit the given control resource from active controllers and release a few control resource of active controllers by reasonably configuring some normal flows under legacy routing mode. Thus, FlexPM can promise ample control resource to improve the recovery efficiency and further effectively map offline switches to active controllers. Simulation results show that FlexPM outperforms existing switch-level solutions by maintaining balanced programmability and increasing the total programmability of recovered offline flows up to 660% under AT&T topology and 590% under Belnet topology. Zehua Guo 0001, Songshi Dou, Wenfei Wu, Yuanqing Xia |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | FlexDATE: Flexible and Disturbance-Aware Traffic Engineering With Reinforcement Learning in Software-Defined NetworksabstractTraffic Engineering (TE) is an important network operation that routes/reroutes flows based on network topology and traffic demands to optimize network performance. Recently, new emerging applications pose challenges to TE with dynamic network conditions, where frequent routing updates are required to maintain good network performance with Software-Defined Networking (SDN). However, flow rerouting operations could lead to considerable Quality of Service (QoS) degradation and service disruption, which is often neglected by existing TE solutions. In this paper, we apply a new QoS metric named network disturbance to measure the negative impact of flow rerouting operations performed by TE. To achieve near-optimal load balancing performance and mitigate network disturbance together in dynamic network scenarios, we propose a flexible and disturbance-aware TE solution called FlexDATE that combines Reinforcement Learning (RL) and Linear Programming (LP). Specifically, FlexDATE leverages RL to intelligently identify flexible numbers of critical flows for each traffic matrix and reroutes these critical flows based on LP optimization to improve network performance with low disturbance. Empowered by a customized actor-critic architecture coupled with Graph Neural Networks (GNNs), FlexDATE can generalize well to unseen traffic scenarios and remain resilient to single link failures. Extensive simulations are conducted on five real-world network topologies to evaluate FlexDATE with real and synthetic traffic traces. The results show that FlexDATE can achieve the performance target (i.e., 90% of optimal performance) in 99% of network scenarios and effectively mitigate the average and maximum network disturbance by up to 9.1% and 38.6%, respectively, compared to state-of-the-art TE solutions. Minghao Ye, Junjie Zhang 0001, Zehua Guo 0001, H. Jonathan Chao |
IEEE/ACM Trans. Netw. | 3 |
| 2022 | ERA: Meeting the Fairness between Sender-driven and Receiver-driven Transmission Protocols in Data Center NetworksabstractThe modern data centers require high throughput and low latency transmission to meet the demands of distributed applications on communication delay. Compared with traditional sender-driven try-and-back-off protocols (e.g., TCP and its variants), receiver-driven protocols (RDPs) achieve the ultra-low transmission latency by reacting to credits or tokens from receivers. However, RDPs face fairness challenges when coexisting with sender-driven protocols (SDPs) in multi-tenant data centers. Their flows barely survive during coexistence with SDP flows since the delicate scheduling of their credits is disrupted and overwhelmed by SDP data packets. To tackle this issue, we propose the Equivalent Rate Adaptor (ERA), a scheme that converts the proactive try-and-back-off mode of SDPs to an RDP-like credit-based reactive mode. ERA leverages the advertised window field in ACK headers at the receiver side to elaborately limit the number of the in-flight packets or bytes in SDPs and thus reduce their impacts on RDPs. Therefore, ERA not only ensures the fairness between two different types of protocols, but also maintains the low latency feature of RDPs. Moreover, ERA is lightweight, flexible, and transparent to tenants by embedding into the prevalent Open vSwitch in the public cloud. The evaluation of both test-bed and NS2 simulation shows that ERA enables SDP flows and RDP flows to maintain good throughput and share the bandwidth fairly, improving the bandwidth stolen by up to 94.29%. Sen Liu 0002, Furong Liang, Zehua Guo 0001, Yang Xu 0010 |
ICDCS | 4 |
| 2022 | ABS: Adaptive Buffer Sizing via Augmented Programmability with Machine LearningabstractProgrammable switches have been proposed in today’s network to enable flexible reconfiguration of devices and reduce time-to-deployment. Buffer sizing, an important factor for network performance, however, has not received enough attention in programmable network. The state-of-the-art buffer sizing solutions usually employ either fixed buffer size or adjust the buffer size heuristically. Without programmability, they suffer from either massive packet drops or large queueing delay in dynamic environment. In this paper, we propose Adaptive Buffer Sizing (ABS), a low-cost and deploy-friendly framework compatible with programmable network. By decoupling the data plane and control plane, ABS-capable switches only need to react to the actions from controller, optimizing network performance in run-time under dynamic traffic. Meanwhile, actions can be programmed by particular Machine Learning (ML) models in the controller to meet different network requirements. In this paper, we address two specific ML models for different scenarios, a reinforcement learning model for relatively stable network with user specific quality requirements, and a supervised learning model for highly dynamic network condition. We implement the ABS framework by integrating the prevalent network simulator NS-2 with ML module. The experiment shows that ABS outperforms state-of-the-art buffer sizing solutions by up to 38.23x under various network environments. Jiaxin Tang, Sen Liu 0002, Yang Xu 0010, Zehua Guo 0001, Junjie Zhang 0001, Peixuan Gao, Yang Chen 0001, Xin Wang 0002, H. Jonathan Chao |
INFOCOM | 4 |
| 2022 | SFP: Service Function Chain Provision on Programmable Switches for Cloud TenantsabstractRecent progress in programmable switches provides opportunities for service function chains (SFCs) provision to cloud tenants, which has the advantage of flexible deployment and high performance. We devise SFP for such SFC provision in the cloud. SFP's data plane installs physical NFs and is virtualized to host logical SFCs from multiple tenants. SFP's control plane uses a relaxed integer programming model to jointly optimize the placement of physical and logical NFs, which can achieve resource efficiency and high tenant traffic processing throughput within efficient execution time. Our prototype and evaluation shows that SFP can significantly offload NFV computation from server to the switch and maximize the switch resource utilization. Hongyi Huang, Wenfei Wu, Yongchao He, Zehua Guo 0001 |
IPDPS | 4 |
| 2022 | TSBS: A Two-Stage Backpressure Scheduling scheme over multihop wireless networks
Chenggang Shan, Yuanqing Xia, Zehua Guo 0001, Jinhui Zhang 0003 |
Ad Hoc Networks | 3 |
| 2022 | Network Coding-based Resilient Routing for Maintaining Data Security and Availability in Software-Defined Networks
Haoran Ni, Zehua Guo 0001, Songshi Dou, Thar Baker |
J. Netw. Comput. Appl. | 2 |
| 2022 | Mitigating Routing Update Overhead for Traffic Engineering by Combining Destination-Based Routing With Reinforcement LearningabstractTraffic Engineering (TE) is a widely-adopted network operation to optimize network performance and resource utilization. Destination-based routing is supported by legacy routers and more readily deployed than flow-based routing, where the forwarding entries could be frequently updated by TE to accommodate traffic dynamics. However, as the network size grows, destination-based TE could render high time complexity when generating and updating many forwarding entries, which may limit the responsiveness of TE and degrade network performance. In this paper, we propose a novel destination-based TE solution called FlexEntry, which leverages emerging Reinforcement Learning (RL) to reduce the time complexity and routing update overhead while achieving good network performance simultaneously. For each traffic matrix, FlexEntry only updates a few forwarding entries calledcritical entriesfor redistributing a small portion of the total traffic to improve network performance. These critical entries are intelligently selected by RL with traffic split ratios optimized by Linear Programming (LP). We find out that the combination of RL and LP is very effective. Our simulation results on six real-world network topologies show that FlexEntry reduces up to 99.3% entry updates on average and generalizes well to unseen traffic matrices with near-optimal load balancing performance. Minghao Ye, Junjie Zhang 0001, Zehua Guo 0001, H. Jonathan Chao |
IEEE J. Sel. Areas Commun. | 4 |
| 2022 | SDN-ESRC: A Secure and Resilient Control Plane for Software-Defined NetworksabstractIn this paper, we propose a resilient control plane based on endogenous security for Software-Defined Networking (SDN) named SDN-ESRC to prevent vulnerability backdoor attacks. SDN-ESRC uses a set of heterogeneous controllers (e.g., RYU, OpenDayLight, ONOS) to compose the control plane and dynamically and adaptively selects several heterogeneous controller instances from the controller set to detect and correct the malicious control messages. The design of SDN-ESRC faces two challenges: (1) increasing network update delay due to multi-controller comparison and (2) maintaining high controllable security. To address the first challenge, SDN-ESRC adopts the master modification mode to reduce the network update delay and identify malicious control messages. To address the second challenge, SDN-ESRC introduces the comparison modification mode to ensure high availability in real time. We propose an evaluation model for SDN-ESRC and theoretically analyze the SDN-ESRC’s endogenous security performance under three typical backdoor attack scenarios. We implement SDN-ESRC in a prototype system and conduct simulations and experiments. The results show that SDN-ESRC can improve the backdoor damage attack security up to 98.3%, the backdoor random attack security up to 99.99%, and the backdoor coordinated attack security up to 82% at the cost of increasing network update delay less than 8.3%. Quan Ren, Zehua Guo 0001, Jiangxing Wu 0001, Tao Hu 0002, Jie Lu 0006 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2022 | SDNShield: NFV-Based Defense Framework Against DDoS Attacks on SDN Control PlaneabstractSoftware-defined networking (SDN) is increasingly popular in today’s information technology industry, but existing SDN control plane is insufficiently scalable to support on-demand, high-frequency flow requests. Weaknesses along SDN control paths can be exploited by malicious third parties to launch distributed denial-of-service (DDoS) attacks against the SDN control plane. Recently proposed solutions only partially solve the problem, by protecting either the SDN network edges or the centralized controller. We propose SDNShield, a solution based on emerging network function virtualization (NFV) technologies, which enforces more comprehensive defense against potential DDoS attacks on SDN control plane. SDNShield incorporates a three-stage overload control scheme. The first stage statistically identifies legitimate flows with low complexity and performance overhead. The second stage further performs in-depth TCP handshake verification to ensure good flows are eventually served. The third stage intellectually salvages the misclassified legitimate flows that are falsely dropped from the first two stages. Prototype tests and real data-driven simulation results show that SDNShield can achieve high resilience against brute-force attacks, and maintain good flow-level service quality at the same time. Kuan-yin Chen, Sen Liu 0002, Yang Xu 0010, Ishant Kumar Siddhrau, Zehua Guo 0001, H. Jonathan Chao |
IEEE/ACM Trans. Netw. | 6 |
| 2022 | Maintaining Control Resiliency and Flow Programmability in Software-Defined WANs During Controller FailuresabstractProviding resilient network control is a critical concern for deploying Software-Defined Networking (SDN) into Wide-Area Networks (WANs). For performance reasons, a Software-Defined WAN is divided into multiple domains controlled by multiple controllers with a logically centralized view. Under controller failures, we need to remap the control of offline switches from failed controllers to other active controllers. Existing solutions have three limitations: (1) the least flow programmability (e.g., the ability to change paths of flows) cannot be maintained; (2) active controllers could be overloaded, interrupting their normal operations; (3) network performance could be degraded because of the increasing controller-switch communication overhead. In this paper, we propose RetroFlow+ to recover the flow programmability and achieve low communication overhead during controller failures. By intelligently configuring a set of selected offline switches working under the legacy routing mode and several active controllers releasing a few control resources, RetroFlow+ enables active controllers to use the minimum control resource to sustain the flow programmability. RetroFlow+ also smartly transfers the control of offline switches with the SDN routing mode to active controllers to minimize the communication overhead from these offline switches to the active controllers. Simulation results show that RetroFlow+ realizes low communication overhead, recovers all offline flows under one and two controller failures, and improves the flow recovery percentage up to 70% under three controller failures, compared with the state-of-the-art solution. Zehua Guo 0001, Songshi Dou, Sen Liu 0002, Wendi Feng, Wenchao Jiang, Yang Xu 0010, Zhi-Li Zhang |
IEEE/ACM Trans. Netw. | 1 |
| 2022 | Enabling Scalable Routing in Software-Defined Networks With Deep Reinforcement Learning on Critical NodesabstractTraditional routing schemes usually use fixed models for routing policies and thus are not good at handling complicated and dynamic traffic, leading to performance degradation (e.g., poor quality of service). Emerging Deep Reinforcement Learning (DRL) coupled with Software-Defined Networking (SDN) provides new opportunities to improve network performance with automatic traffic analysis and policy generation. However, existing DRL-based routing solutions usually rely on all node information to make routing decisions for the network and hence are both hard to converge in large networks and vulnerable to topology changes. In this paper, we propose ScaleDeep, a scalable DRL-based routing scheme for SDN, which improves the routing performance and is resilient to topology changes. Essentially, ScaleDeep takes advantage of partial control on network nodes and DRL. We select a set of critical nodes from a network as driver nodes, which can simulate the entire network operation, based on the control theory. By observing the traffic variation on the driver nodes, DRL dynamically adjusts some link weights for a weighted shortest path algorithm to change the routing paths and improve the routing performance. Limiting the control on driver nodes improves the convergence ability of DRL and reduces the dependency of the DRL agent on the fixed network topology. To validate the performance of ScaleDeep, we conduct packet-level simulations on different topologies. The results show that ScaleDeep outperforms existing DRL-based schemes by reducing the average flow completion time by up to 36% and exhibiting better robustness against minor topology changes. Penghao Sun, Zehua Guo 0001, Junfei Li, Yang Xu 0010, Julong Lan, Yuxiang Hu 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2021 | Exploring the Impact of Attacks on Ring AllReduceabstractDistributed Machine Learning (DML) is widely used to accelerate the training of the deep learning model. In DML, Parameter-Server (PS) and Ring AllReduce are two typical architectures. Recently, observing that many works address the security problem in PS, whose performance can be greatly degraded by malicious participation during the training process. However, the robustness of Ring AllReduce, which can solve the communication bandwidth problem in PS, to the malicious participant is still unknown. In this paper, we design a series of experiments to explore the security problem in Ring AllReduce, and reveal it can also suffer from the malicious participant. Zehua Guo 0001, Sen Liu 0002 |
APNet | 3 |
| 2021 | ProgrammabilityMedic: Predictable Path Programmability Recovery under Multiple Controller Failures in SD-WANsabstractSoftware-Defined Networking (SDN) promises good network performance in Wide Area Networks (WANs) with the logically centralized control using physically distributed controllers. In Software-Defined WANs (SD-WANs), maintaining path programmability, which enables flexible path change on flows, is crucial for maintaining network performance under traffic variation. However, when controllers fail, existing solutions are essentially coarse-grained switch-controller mapping solutions and only recover the path programmability of a limited number of offline flows, which traverse offline switches controlled by failed controllers. In this paper, we propose ProgrammabilityMedic (PM) to provide predictable path programmability recovery under controller failures in SD-WANs. The key idea of PM is to approximately realize flow-controller mappings using hybrid SDN/legacy routing supported by high-end commercial SDN switches. Using the hybrid routing, we can recover programmability by fine-grainedly selecting a routing mode for each offline flow at each offline switch to fit the given control resource from active controllers. Thus, PM can effectively map offline switches to active controllers to improve recovery efficiency. Simulation results show that PM outperforms existing switch-level solutions by maintaining balanced programmability and increasing the total programmability of recovered offline flows up to 315% under two controller failures and 340% under three controller failures. Songshi Dou, Zehua Guo 0001, Yuanqing Xia |
ICDCS | 2 |
| 2021 | Poster: Enabling Fast Forwarding in Hybrid Software-Defined NetworksabstractEmerging Software-Defined Networking (SDN) technique brings new opportunities to improve network performance. Some SDN-enabled programmable switches are deployed in legacy networks, and thus legacy and programmable switches could coexist, generating hybrid SDNs. In this paper, we study the node upgrade for layer-2 hybrid SDN and propose Shortcutter to accelerate the transmission. Preliminary results show that the proposed Shortcutter can reduce the forwarding path’s length 7% on average, compared with baseline solutions. Yijun Sun, Zehua Guo 0001, Songshi Dou, Junjie Zhang 0001, Xiang Ouyang |
ICNP | 2 |
| 2021 | Federated Traffic Engineering with Supervised Learning in Multi-region NetworksabstractNetwork operators usually adopt Traffic Engineering (TE) to configure the routing in their networks to achieve good load balancing performance and high resource utilization. While centralized TE can effectively improve network performance with a global view of the network, distributed TE has been considered as an alternative to manage large-scale networks that are usually partitioned into multiple regions. However, it is challenging for distributed TE to reach a global optimal performance since each region can make its local routing decisions only based on partially observed network states. In this paper, we propose a novel distributed TE scheme called FedTe, which leverages supervised learning coupled with a collaborative approach to improve the overall load balancing performance for multi-region networks. FedTe learns from the global optimal routing strategy in a centralized offline manner and predicts the optimal distribution of cross-region traffic among different regions through distributed deployment in real time. The predicted cross-region traffic distribution is integrated with measured local traffic to construct each region’s optimal regional traffic matrix, which is used to perform intra-region TE optimization. FedTe can also handle dynamic traffic variation and link failures with a 2-layer hierarchical graph neural network architecture. To validate the effectiveness of the proposed scheme, we evaluate FedTe with two real-world network topologies and a large-scale synthetic topology. Extensive evaluation results show that FedTe can achieve near-optimal load balancing performance and outperform state-of-the-art distributed TE approaches by up to 28.9% on average. Minghao Ye, Junjie Zhang 0001, Zehua Guo 0001, H. Jonathan Chao |
ICNP | 3 |
| 2021 | Optimizing Flow Completion Time via Adaptive Buffer Management in Data Center NetworksabstractThe traffic of modern data centers exhibits long-tail distribution, in which massive delay-sensitive short flows and a small number of bandwidth-hungry long flows co-exist. These two types of flows could share same bottleneck links in the data center networks but request different or even opposite network requirements. Existing solutions try to realize a trade-off between the requirements of different flows by either prioritizing short flows or limiting the buffer used by long flows at switches or end-hosts. However, they do not consider the dynamic traffic change and suffer from performance degradation, resulted from severe queueing delay and massive packet drops for short flows under current First-In-First- Out (FIFO) queueing mechanism. In this paper, we propose a novel buffer management scheme at switches, called Cut-in Queue (CQ), to achieve both low latency for short flows and high throughput for long flows. Based on network status in real time, CQ prioritizes short flows by dynamically cutting the short flows’ packets into the head of long flows or evicting some enqueued long flows’ packets and enables high throughput for long flows in most of the cases. Evaluation of both DPDK testbed and NS2 simulations show that CQ outperforms state-of-the-art buffer management schemes by reducing flow completion time by up to 73%. Sen Liu 0002, Zehua Guo 0001, Yi Wang 0004, Mohamed Adel Serhani, Yang Xu 0010 |
ICPP | 3 |
| 2021 | Video Quality and Popularity-aware Video Caching in Content Delivery NetworksabstractContent Delivery Network (CDN) is a popular service to accelerate object transmission by dynamically caching popular objects at cache points near users. Existing video caching schemes for CDN do not consider some important components of Quality of Experience (QoE). In this paper, we jointly consider video quality and popularity to design a new QoE metric called Video Hit Experience (VHE) and propose an efficient video caching algorithm named Hit ExpeRience-based videO caching (HERO) to improve VHE. Preliminary results show that HERO outperforms existing solutions. Yijun Sun, Zehua Guo 0001, Songshi Dou, Yuanqing Xia |
ICWS | 2 |
| 2021 | DATE: Disturbance-Aware Traffic Engineering with Reinforcement Learning in Software-Defined NetworksabstractTraffic Engineering (TE) has been applied to optimize network performance by routing/rerouting flows based on traffic loads and network topologies. To cope with network dynamics from emerging applications, it is essential to reroute flows more frequently than today’s TE to maintain network performance. However, existing TE solutions may introduce considerable Quality of Service (QoS) degradation and service disruption since they do not take the potential negative impact of flow rerouting into account. In this paper, we apply a new QoS metric named network disturbance to gauge the impact of flow rerouting while optimizing network load balancing in backbone networks. To employ this metric in TE design, we propose a disturbance-aware TE called DATE, which uses Reinforcement Learning (RL) to intelligently select some critical flows between nodes for each traffic matrix and reroute them using Linear Programming (LP) to jointly optimize network performance and disturbance. DATE is equipped with a customized actor-critic architecture and Graph Neural Networks (GNNs) to handle dynamic traffic and single link failures. Extensive evaluations show that DATE can outperform state-of-the-art TE methods with close-to-optimal load balancing performance while effectively mitigating the 99th percentile network disturbance by up to 31.6%. Minghao Ye, Junjie Zhang 0001, Zehua Guo 0001, H. Jonathan Chao |
IWQoS | 3 |
| 2021 | Matchmaker: Maintaining network programmability for Software-Defined WANs under multiple controller failures
Songshi Dou, Guochun Miao, Zehua Guo 0001, Weiran Wu, Yuanqing Xia |
Comput. Networks | 3 |
| 2021 | ScaleDRL: A Scalable Deep Reinforcement Learning Approach for Traffic Engineering in SDN with Pinning Control
Penghao Sun, Zehua Guo 0001, Julong Lan, Junfei Li, Yuxiang Hu 0001, Thar Baker |
Comput. Networks | 2 |
| 2021 | An adaptive defense mechanism to prevent advanced persistent threatsabstractThe expansion of information technology infrastructure is encountered with Advanced Persistent Threats (APTs), which can launch data destruction, disclosure, modification, and/or Denial of Service attacks by drawing upon vulnerabilities of software and hardware. Moving Target Defense (MTD) is a promising risk mitigation technique that replies to APTs via implementing randomisation and dynamic strategies on compromised assets. However, some MTD techniques adopt the blind random mutation, which causes greater performance overhead and worse defense utility. In this paper, we formulate the cyber-attack and defense as a dynamic partially observable Markov process based on dynamic Bayesian inference. Then we develop an Inference-Based Adaptive Attack Tolerance (IBAAT) system , which includes two stages. In the first stage, a forward–backward algorithm with a time window is employed to perform a security risk assessment. To select the defense strategy, in the second stage, the attack and defense process is modelled as a two-player general-sum Markov game and the optimal defense strategy is acquired by quantitative analysis based on the first stage. The evaluation shows that the proposed algorithm has about 10% security utility improvement compared to the state-of-the-art. Yi-xi Xie, Li-xin Ji, Ling-shu Li, Zehua Guo 0001, Thar Baker |
Connect. Sci. | 4 |
| 2021 | Enabling Technologies for Energy Cloud
Thar Baker, Zehua Guo 0001, Ali Ismail Awad, Shangguang Wang, Benjamin C. M. Fung |
J. Parallel Distributed Comput. | 2 |
| 2021 | HybridFlow: Achieving Load Balancing in Software-Defined WANs With Scalable RoutingabstractThe scalability issue hinders the deployment of Software-Defined Networking (SDN) in the Wide Area Networks (WANs). Existing solutions have two issues: (1) network performance relies on complicated controller synchronization, which increases the complexity of network control; (2) fine-grained flow processing enables flexible flow control at the cost of high processing load on the controllers and high flow table occupancy on switches. In this paper, we propose a scalable routing solution named HybridFlow, which achieves a good load balancing performance using a single controller with low control overhead (i.e., flow routing and rerouting overhead). HybridFlow mainly employs two techniques: hybrid routing and crucial flow rerouting. Hybrid routing enabled by commercial SDN switches gives us opportunities to reduce the processing load of the controller by routing flows with the hybrid OpenFlow/OSPF mode. Thus, the majority of flows can be routed by OSPF without involving the controller. Crucial flow rerouting realizes load balancing by dynamically identifying crucial flows based on a new metric called Variation Slope and rerouting these flows with the hybrid OpenFlow/OSPF mode. The simulation based on the real traffic traces and network typologies shows that compared with the optimal solution, HybridFlow can achieve 87% of the optimal load balancing performance by rerouting 36% less flows on average. Zehua Guo 0001, Songshi Dou, Yi Wang 0004, Sen Liu 0002, Wendi Feng, Yang Xu 0010 |
IEEE Trans. Commun. | 1 |
| 2021 | AggreFlow: Achieving Power Efficiency, Load Balancing, and Quality of Service in Data Center NetworksabstractPower-efficient Data Center Networks (DCNs) have been proposed to save power of DCNs using OpenFlow. In these DCNs, the OpenFlow controller adaptively turns on/off links and OpenFlow switches to form a minimum-power subnet that satisfies the traffic demand. As the subnet changes, flows are dynamically routed and rerouted to the routes composed of active switches and links. However, existing flow scheduling schemes could cause undesired results: (1) power inefficiency: due to unbalanced traffic allocation on active routes, extra switches and links may be activated to cater to bursty traffic surges on congested routes, and (2) Quality of Service (QoS) fluctuation: because of the limited flow entry processing ability, switches may not be able to timely install/delete/update flow entries to properly route/reroute flows. In this paper, we propose AggreFlow, a dynamic flow scheduling scheme that achieves power efficiency and QoS improvement using three techniques: Flow-set Routing, Lazy Rerouting, and Adaptive Rerouting. Flow-set Routing achieves load balancing with a small number of flow entry operations by routing flows in a coarse-grained flow-set fashion. Lazy Rerouting spreads rerouting operations over a relatively long period of time, reducing the burstiness of entry operation on switches. Adaptive Rerouting selectively reroutes flow-sets to maintain load balancing. We built an NS3 based fat-tree network simulation platform to evaluate AggreFlow's performance. The simulation results show that AggreFlow reduces power consumption by about 18%, yet achieving load balancing and improved QoS (low packet loss rate and reducing the number of processing entries for flow scheduling by 98%), compared with baseline schemes. Zehua Guo 0001, Yang Xu 0010, Ya-Feng Liu, Sen Liu 0002, H. Jonathan Chao, Zhi-Li Zhang, Yuanqing Xia |
IEEE/ACM Trans. Netw. | 1 |
| 2020 | QOS-Aware Flow Control for Power-Efficient Data Center Networks with Deep Reinforcement LearningabstractReducing the power consumption and maintaining the Flow Completion Time (FCT) for the Quality of Service (QoS) of applications in Data Center Networks (DCNs) are two major concerns for data center operators. However, existing works either fail in guaranteeing the QoS due to the neglect of the FCT constraints or achieve a less satisfying power efficiency. In this paper, we propose SmartFCT, which employs Software-Defined Networking (SDN) coupled with the Deep Reinforcement Learning (DRL) to improve the power efficiency of DCNs and guarantee the FCT. The DRL agent can generate a dynamic policy to consolidate traffic flows into fewer active switches in the DCN for power efficiency, and the policy also leaves different margins in different active links and switches to avoid FCT violation of unexpected short bursts of flows. Simulation results show that with similar FCT guarantee, SmartFCT can save 8% more of the power consumption compared to the state-of-the-art solutions. Penghao Sun, Zehua Guo 0001, Sen Liu 0002, Julong Lan, Yuxiang Hu 0001 |
ICASSP | 2 |
| 2020 | Improving the Scalability of Deep Reinforcement Learning-Based Routing with Control on Partial NodesabstractMachine Learning (ML)-based routing optimization has been proposed to optimize the performance of flow routing for future networks, such as Software-Defined Networks (SDNs). However, existing studies are either hard to converge for large networks or vulnerable to topology changes. In this paper, we propose SINET, a scalable and intelligent network control framework for routing optimization. To improve the robustness and scalability, SINET selects several critical routing nodes to be directly controlled by a Deep Reinforcement Learning (DRL) agent, which dynamically generates routing policy to optimize network performance. Simulation results show that SINET can reduce the average flow completion time by at least 32% for a network with 82 nodes and exhibit better robustness against minor topology changes, compared to other DRL-based schemes. Penghao Sun, Julong Lan, Zehua Guo 0001, Yang Xu 0010, Yuxiang Hu 0001 |
ICASSP | 3 |
| 2020 | BAGUETTE: Towards a Secure and Cost-effective Switch Upgrade in Hybrid Software-Defined NetworksabstractSoftware-Defined Networking (SDN), providing flexible controlling and monitoring mechanisms that simplifies network management, is becoming prevalent in recent years. However, replacing all legacy network devices with SDN-capable devices is cost-prohibitive. One practical approach for the SDN deployment is to incrementally upgrade a few legacy devices to SDN devices. The network, which consists of legacy and SDN devices, is called a hybrid SDN. Existing hybrid SDN deployment schemes do not consider the security impact of device deployment. They use the same type of devices to upgrade, and upgraded devices could be compromised if an attacker controls one SDN device by leveraging its vulnerabilities.In this paper, we consider this security issue in the hybrid SDN deployment and present the Secure and Cost-effective Switch Upgrade (SCESU) problem. The SCESU problem aims to upgrade a few network devices to satisfy the security requirement by using multiple SDN switch types with a minimal upgrade cost. The complexity of the SCESU problem comes from common vulnerabilities shared among different types of SDN devices and attack propagations among network nodes. To efficiently solve the problem, we propose the BAGUETTE algorithm to judiciously choose and upgrade critical legacy switches with selected SDN devices. Simulation results show that BAGUETTE achieves up to about 92.1% security enhancement compared with legacy network and reduces to 11.1% cost of the securest deployment. Wendi Feng, Zehua Guo 0001, Chuanchang Liu, Yueming Zheng, Meng Wang 0018, Bo Cheng 0001, Junliang Chen 0001 |
ICC | 2 |
| 2020 | DeepMigration: Flow Migration for NFV with Graph-based Deep Reinforcement LearningabstractNetwork Function Virtualization (NFV) enables flexible deployment of network services as applications. Network operators expect to use a limited number of Network Function (NF) instances to handle the fluctuating traffic load and provide network services. However, it is a big challenge to guarantee the Quality of Service (QoS) under the unpredictable network traffic while minimizing the processing resources. One typical solution is to realize NF scale-out, scale-in and load balancing by elastically migrating the related traffic flows with SoftwareDefined Networking (SDN). However, it is difficult to optimally migrate flows since many real-time statuses of NF instances should be considered to make accurate decisions. In this paper, we propose DeepMigration to solve the problem by efficiently and dynamically migrating traffic flows among different NF instances. DeepMigration is a Deep Reinforcement Learning (DRL)-based solution coupled with Graph Neural Network (GNN). By taking advantages of the graph-based relationship deduction ability from our customized GNN and the self-evolution ability from the experience training of DRL, DeepMigration can accurately model the cost (e.g., migration latency) and the benefit (e.g., reducing the number of NF instances) of flow migration among different NF instances and generate dynamic and effective flow migration policies to improve the QoS. Experiment results show that DeepMigration requires less migration cost and saves up to 71.6{%} of the computation time than existing solutions. Penghao Sun, Julong Lan, Zehua Guo 0001, Di Zhang 0002, Xianfu Chen, Yuxiang Hu 0001, Zhi Liu 0002 |
ICC | 3 |
| 2020 | Poster: Maintaining Training Efficiency and Accuracy for Edge-assisted Online Federated Learning with ABSabstractThis paper proposes Adaptive Batch Sizing (ABS) for online federated learning. ABS is an iteration process-efficient solution that adaptively adjusts batch size of the training process at edge nodes. Preliminary results show that ABS maintains training efficiency and accuracy, compared with existing iteration round-efficient solutions. Zehua Guo 0001, Sen Liu 0002, Yuanqing Xia |
ICNP | 2 |
| 2020 | DeepWeave: Accelerating Job Completion Time with Deep Reinforcement Learning-based Coflow SchedulingabstractTo improve the processing efficiency of jobs in distributed computing, the concept of coflow is proposed. A coflow is a collection of flows that are semantically correlated in a multi-stage computation task. A job consists of multiple coflows and can be usually formulated as a Directed-Acyclic Graph (DAG). A proper scheduling of coflows can significantly reduce the completion time of jobs in distributed computing. However, this scheduling problem is proved to be NP-hard. Different from existing schemes that use hand-crafted heuristic algorithms to solve this problem, in this paper, we propose a Deep Reinforcement Learning (DRL) framework named DeepWeave to generate coflow scheduling policies. To improve the inter-coflow scheduling ability in the job DAG, DeepWeave employs a Graph Neural Network (GNN) to process the DAG information. DeepWeave learns from the history workload trace to train the neural networks of the DRL agent and encodes the scheduling policy in the neural networks, which make coflow scheduling decisions without expert knowledge or a pre-assumed model. The proposed scheme is evaluated with a simulator using real-life traces. Simulation results show that DeepWeave completes jobs at least 1.7X faster than the state-of-the-art solutions. Penghao Sun, Zehua Guo 0001, Junfei Li, Julong Lan, Yuxiang Hu 0001 |
IJCAI | 2 |
| 2020 | Improving the Path Programmability for Software-Defined WANs under Multiple Controller FailuresabstractEnabling path programmability is an essential feature of Software-Defined Networking (SDN). During controller failures in Software-Defined Wide Area Networks (SD-WANs), a resilient design should maintain path programmability for offline flows, which were controlled by the failed controllers. Existing solutions can only partially recover the path programmability rooted in two problems: (1) the implicit preferable recovering flows with long paths and (2) the sub-optimal remapping strategy in the coarse-grained switch level. In this paper, we propose Programmability Guardian to improve the path programmability of offline flows while maintaining low communication overhead. These goals are achieved through the fine-grained flow-level mappings enabled by existing SDN techniques. Programmability Guardian configures the flow-controller mappings to recover offline flows with a similar path programmability, maximize the total programmability of the offline flows, and minimize the total communication overhead for controlling these recovered flows. Simulation results of different controller failure scenarios show that Programmability Guardian recovers all offline flows with a balanced path programmability, improves the total programmability of the recovered flows up to 68%, and reduces the communication overhead up to 83%, compared with the baseline algorithm. Zehua Guo 0001, Songshi Dou, Wenchao Jiang |
IWQoS | 1 |
| 2020 | SmartFCT: Improving power-efficiency for data center networks with deep reinforcement learning
Penghao Sun, Zehua Guo 0001, Sen Liu 0002, Julong Lan, Yuxiang Hu 0001 |
Comput. Networks | 2 |
| 2020 | MARVEL: Enabling controller load balancing in software-defined networks with multi-agent reinforcement learning
Penghao Sun, Zehua Guo 0001, Gang Wang 0014, Julong Lan, Yuxiang Hu 0001 |
Comput. Networks | 2 |
| 2020 | Efficient flow migration for NFV with Graph-aware deep reinforcement learning
Penghao Sun, Julong Lan, Junfei Li, Zehua Guo 0001, Tao Hu 0002 |
Comput. Networks | 4 |
| 2020 | Mitigating malicious packets attack via vulnerability-aware heterogeneous network devices assignment
Jianjian Ai, Hongchang Chen, Zehua Guo 0001, Thar Baker |
Future Gener. Comput. Syst. | 3 |
| 2020 | MobiGyges: A mobile hidden volume for preventing data loss, improving storage utilization, and avoiding device reboot
Wendi Feng, Chuanchang Liu, Zehua Guo 0001, Thar Baker, Gang Wang 0014, Meng Wang 0018, Bo Cheng 0001, Junliang Chen 0001 |
Future Gener. Comput. Syst. | 3 |
| 2020 | Exploring the role of paths for dynamic switch assignment in software-defined networks
Zehua Guo 0001, Shaojun Zhang, Wendi Feng, Weichao Wu, Julong Lan |
Future Gener. Comput. Syst. | 1 |
| 2020 | CLOSURE: A cloud scientific workflow scheduling algorithm based on attack-defense game model
Zehua Guo 0001, Thar Baker, Wenyan Liu 0005 |
Future Gener. Comput. Syst. | 3 |
| 2020 | Protecting scientific workflows in clouds with an intrusion tolerant systemabstractWith the development of cloud computing technology, more and more scientific workflows are delivered to cloud platforms to complete. However, there are many threats in clouds due to the multi‐tenant coexistence. In order to protect scientific workflows in clouds, the authors propose an intrusion tolerant scientific workflow system. In this system, the task executors containing multiple virtual machines are used for workflow sub‐task execution to enhance reliability. Then lagged decision mechanism is presented to ensure uninterrupted workflow execution while checking the intermediate data, and assessing the confidence of these data. Inspired by moving target defence, they propose a dynamic task scheduling strategy based on resource circulation to periodically generate and recycle task executors, keeping the clean state of the workflow execution environment. Furthermore, temporary workflow intermediate data backup mechanism is presented, the stored intermediate data can be used for the re‐execution of workflow sub‐tasks with low confidence. Experiments are conducted in both the actual test environment based on OpenStack and the simulated test environment based on WorkflowSim toolkit. Experimental results demonstrate that the proposed system can effectively enhance intrusion tolerance of scientific workflows. Zehua Guo 0001, Wenyan Liu 0005, Chao Yang 0012 |
IET Inf. Secur. | 3 |
| 2020 | COMITMENT: A Fog Computing Trust Management Approach
Mohammed Al-Khafajiy, Thar Baker, Muhammad Asim 0001, Zehua Guo 0001, Rajiv Ranjan 0001, Antonella Longo, Deepak Puthal, Mark Taylor 0005 |
J. Parallel Distributed Comput. | 4 |
| 2020 | CFR-RL: Traffic Engineering With Reinforcement Learning in SDNabstractTraditional Traffic Engineering (TE) solutions can achieve the optimal or near-optimal performance by rerouting as many flows as possible. However, they do not usually consider the negative impact, such as packet out of order, when frequently rerouting flows in the network. To mitigate the impact of network disturbance, one promising TE solution is forwarding the majority of traffic flows using Equal-Cost Multi-Path (ECMP) and selectively rerouting a few critical flows using Software-Defined Networking (SDN) to balance link utilization of the network. However, critical flow rerouting is not trivial because the solution space for critical flow selection is enormous. Moreover, it is impossible to design a heuristic algorithm for this problem based on fixed and simple rules, since rule-based heuristics are unable to adapt to the changes of the traffic matrix and network dynamics. In this paper, we propose CFR-RL (Critical Flow Rerouting-Reinforcement Learning), a Reinforcement Learning-based scheme that learns a policy to select critical flows for each given traffic matrix automatically. CFR-RL then reroutes these selected critical flows to balance link utilization of the network by formulating and solving a simple Linear Programming (LP) problem. Extensive evaluations show that CFR-RL achieves near-optimal performance by rerouting only 10%-21.3% of total traffic. Junjie Zhang 0001, Minghao Ye, Zehua Guo 0001, Chen-Yu Yen, H. Jonathan Chao |
IEEE J. Sel. Areas Commun. | 3 |
| 2019 | Improving Resiliency of Software-Defined Networks with Network Coding-based Multipath RoutingabstractTraditional network routing protocol exhibits high statics and singleness, which provide significant advantages for the attacker. There are two kinds of attacks on the network: active attacks and passive attacks. Existing solutions for those attacks are based on replication or detection, which can deal with active attacks; but are helpless to passive attacks. In this paper, we adopt the theory of network coding to fragment the data in the Software-Defined Networks and propose a network coding-based resilient multipath routing scheme. First, we present a new metric named expected eavesdropping ratio to measure the resilience in the presence of passive attacks. Then, we formulate the network coding-based resilient multipath routing problem as an integer-programming optimization problem by using expected eavesdropping ratio. Since the problem is NP-hard, we design a Simulated Annealing-based algorithm to efficiently solve the problem. The simulation results demonstrate that the proposed algorithms improve the defense performance against passive attacks by about 20% when compared with baseline algorithms. Jianjian Ai, Hongchang Chen, Zehua Guo 0001, Thar Baker |
ISCC | 3 |
| 2019 | Data Loss Prevention and Storage Utilization Improvement of the Hidden Volume on Mobile DevicesabstractSensitive data protection is vital for mobile users. An effective way is to store sensitive data in the hidden volume of mobile devices with Plausibly Deniable Encryption (PDE) systems. Typically, PDE creates a hidden volume inside the outer volume. However, existing PDE systems could lose data, due to overriding the hidden volume, and significantly waste physical storage, because of the fixedly reserved area for the hidden volume. In this paper, we present MobiGyges to solve the above problems. MobiGyges leverages the Thin Pool to coordinate the storage allocation for both the outer volume and the hidden volume which avoids data override and prevents data loss. Moreover, it improves the storage efficiency by virtualizing the total physical storage into small storage blocks and utilizing the blocks for the outer volume and the hidden volume, which eliminates the reserved area. We implement MobiGyges prototype on Google Nexus 6P with LineageOS 13 operating system. Experimental results show that MobiGyges avoids data loss and improves storage utilization up to 31% compared with existing works. Wendi Feng, Chuanchang Liu, Zehua Guo 0001, Thar Baker, Bo Cheng 0001, Junliang Chen 0001 |
ISCC | 3 |
| 2019 | RetroFlow: maintaining control resiliency and flow programmability for software-defined WANsabstractProviding resilient network control is a critical concern for deploying Software-Defined Networking (SDN) into Wide-Area Networks (WANs). For performance reasons, a Software-Defined WAN is divided into multiple domains controlled by multiple controllers with a logically centralized view. Under controller failures, we need to remap the control of offline switches from failed controllers to other active controllers. Existing solutions could either overload active controllers to interrupt their normal operations or degrade network performance because of increasing the controller-switch communication overhead. In this paper, we propose RetroFlow to achieve low communication overhead without interrupting the normal processing of active controllers during controller failures. By intelligently configuring a set of selected offline switches working under the legacy routing mode, RetroFlow relieves the active controllers from controlling the selected offline switches while maintaining the flow programmability (e.g., the ability to change paths of flows) of SDN. RetroFlow also smartly transfers the control of offline switches with the SDN routing mode to active controllers to minimize the communication overhead from these offline switches to the active controllers. Simulation results show that compared with the baseline algorithm, RetroFlow can reduce the communication overhead up to 52.6% during a moderate controller failure by recovering 100% flows from offline switches and can reduce the communication overhead up to 61.2% during a serious controller failure by setting to recover 90% of flows from offline switches. Zehua Guo 0001, Wendi Feng, Sen Liu 0002, Wenchao Jiang, Yang Xu 0010, Zhi-Li Zhang |
IWQoS | 1 |
| 2019 | Chunk-level request-grant-transfer mode for QoE-sensitive video delivery in CDNabstractRemote Direct Memory Access (RDMA) can be deployed in Content Delivery Networks (CDN) Points of Presence (PoPs) to avoid the high CPU overheads caused by traditional TCP/IP stacks. However, RDMA cannot surmount the drawbacks of the window-based conservative of TCP and is insensitive to Quality of Experience (QoE). Moreover, the requirement of lossless networks hinders the widespread application of RDMA. In this paper, we introduce the parallel multipoint-to-multipoint Request-Grant-Transfer (RGT) mode into RDMA to solve the aforementioned problems. Compared with traditional RGT mode, our scheme supports parallel Dynamic Adaptive Streaming over HTTP (DASH) chunk delivery, thereby improving throughput and reducing initial delays. We differentiate the importance of DASH chunks according to QoE-related properties. In this way, we reduce the response time of specific DASH chunks. We provide an efficient approach to select the optimal number of requests for partially traversing pending requests to reduce the overheads of Request stages. We perform comprehensive experiments to demonstrate that our scheme improves the throughput of CDN PoPs and enhances client QoE. Gengbiao Shen, Qing Li 0006, Yong Jiang 0001, Richard O. Sinnott, Dong Lin, Zehua Guo 0001, Yi Wang 0004 |
IWQoS | 6 |
| 2019 | Dynamic slave controller assignment for enhancing control plane robustness in software-defined networks
Tao Hu 0002, Peng Yi 0003, Zehua Guo 0001, Julong Lan |
Future Gener. Comput. Syst. | 3 |
| 2019 | Joint Switch Upgrade and Controller Deployment in Hybrid Software-Defined NetworksabstractTo improve traffic management ability, Internet Service Providers (ISPs) are gradually upgrading legacy network devices to programmable devices that support Software-Defined Networking (SDN). The coexistence of legacy and SDN devices gives rise to a hybrid SDN. Existing hybrid SDNs do not consider the potential performance issues introduced by a centralized SDN controller: flow requests processed by a highly loaded controller may experience long-tail processing delay; inappropriate multi-controller deployment could increase the propagation delay of flow requests. In this paper, we propose to jointly consider the deployment of SDN switches and their controllers for hybrid SDNs. We formulate the joint problem as an optimization problem that maximizes the number of flows that can be controlled and managed by the SDN and minimizes the propagation delay of flow requests between SDN controllers and switches under a given upgrade budget constraint. We show this problem is NP-hard. To efficiently solve the problem, we propose some techniques (e.g., strengthening the constraints and adding additional valid inequalities) to accelerate the global optimization solver for solving the problem for small networks and an efficient heuristic algorithm for solving it for large networks. The simulation results from real network topologies illustrate the effectiveness of the proposed techniques and show that our proposed heuristic algorithm uses a small number of controllers to manage a high amount of flows with good performance. Zehua Guo 0001, Ya-Feng Liu, Yang Xu 0010, Zhi-Li Zhang |
IEEE J. Sel. Areas Commun. | 1 |
| 2019 | RAPID: Avoiding TCP Incast Throughput Collapse in Public Clouds With Intelligent Packet DiscardingabstractMany applications in public clouds require a high fan-in, many-to-one type of data communication (known as TCP incast) in modern Data Center Networks (DCNs). Such communication could cause severe incast congestion in switches and result in TCP throughput collapse, substantially degrading the application performance. The root cause of throughput collapse is the Retransmission Timeouts (RTO) due to packet losses in congested switches. Tenants in public clouds can opt to use a variety of TCP versions. However, the existing solutions rely on modifications of TCP protocols and specific techniques from switches, and thus these existing solutions are not always feasible for public clouds. In this paper, we are inspired by the emerging virtualization and network softwarization technologies to develop a novel scheme called Retransmission timeout Avoidance by Packet Intelligent Discarding (RAPID) using software switches. RAPID considers the number of packets of each incast flow, buffered in the switch to selectively discard some packets, and ensures that the Fast Retransmission/Fast Recovery rather than RTO is invoked at the sender(s) in response to packet loss. Thus, the long idle period of a timeout and the throughput drop are avoided. We prove that, given a predetermined minimum switch buffer space, dedicated to the incast application, RAPID can prevent RTO in all the incast senders. We also present a low-complexity heuristic version of RAPID named RAPID-ED, which combines the principles of RAPID and early detection and is extremely easy to implement on today's software switches. We evaluate the two proposed schemes in a data center network testbed built on NS-3 simulator. The simulation results confirm the theoretical expectation, and show that the RAPID and RAPID-ED perform very well to prevent RTO of TCP incast flows and hence the throughput collapse. Compared with other incast solutions, RAPID and RAPID-ED do not modify TCP protocols and therefore are more suitable in public clouds. Yang Xu 0010, Shikhar Shukla, Zehua Guo 0001, Sen Liu 0002, Adrian Sai-Wah Tam, Kang Xi, H. Jonathan Chao |
IEEE J. Sel. Areas Commun. | 3 |
| 2018 | Adaptive Slave Controller Assignment for Fault-Tolerant Control Plane in Software-Defined NetworkingabstractMulti-controller is a promising control plane solution for the large-scale Software-Defined Networks (SDN). Some existing works (e.g., OpenFlow 1.2) propose to use backup controllers named slave controllers to achieve fault- tolerance in the control plane. In this paper, we identify the unreasonable slave controller assignment could cause the controller chain failure and eventually crash the entire network. We consider some important factors for designing fault-tolerant control plane and formulating Slave Controller Assignment (SCA) problem. SCA is an NP-complete problem, and we solve it with Adaptive Slave Controller Assignment (ASCA) scheme, which adaptively assigns slave controller according to load variance difference. The numerical results validate the efficiency of ASCA. Tao Hu 0002, Zehua Guo 0001, Julong Lan |
ICC | 2 |
| 2018 | Intelligent path control for energy-saving in hybrid SDN networks
Xuya Jia, Yong Jiang 0001, Zehua Guo 0001, Gengbiao Shen, Lei Wang 0071 |
Comput. Networks | 3 |
| 2018 | Balancing flow table occupancy and link utilization in software-defined networks
Zehua Guo 0001, Yang Xu 0010, Ruoyan Liu, Andrey Gushchin, Kuan-yin Chen, Anwar Elwalid, H. Jonathan Chao |
Future Gener. Comput. Syst. | 1 |
| 2018 | BWManager: Mitigating Denial of Service Attacks in Software-Defined Networks Through Bandwidth PredictionabstractSoftware-defined networking (SDN) has emerged as a new networking paradigm that can provide fine-grained network management service. Since the SDN controller makes control decision for the network, it becomes the main target of denial of service (DoS) attacks. In this paper, we propose BWManager to mitigateE which mainly consists mitigate the DoS attacks on the SDN controller with BWManager that mainly consists of four key components: 1) simplified DoS detection module; 2) forecasting engine; 3) priority manager; and 4) scheduler. The simplified DoS detection module calculates a comprehensive judgment score for each switch, which indicates the attacking severity of each switch and is used to decide time slice allocation of the controller. The forecasting engine is the basis of the controller scheduling method and forecasts the bandwidth consumption of users to determine the users' trust values. The trust values are used by the priority manager to manage multiple buffer queues with different priorities for the users. The scheduler protects the controller and the normal users under DoS attacks by running a weighted Round-Robin algorithm to process flow requests in different priority queues. We evaluate the performance and overhead of BWManager in both hardware and software OpenFlow environments. The results demonstrate that BWManager is effective with a limited overhead. Tao Wang 0018, Zehua Guo 0001, Hongchang Chen |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2017 | STAR: Preventing flow-table overflow in software-defined networks
Zehua Guo 0001, Ruoyan Liu, Yang Xu 0010, Andrey Gushchin, Anwar Elwalid, H. Jonathan Chao |
Comput. Networks | 1 |
| 2017 | A low overhead flow-holding algorithm in software-defined networks
Xuya Jia, Qing Li 0006, Yong Jiang 0001, Zehua Guo 0001 |
Comput. Networks | 4 |
| 2016 | Dynamic flow scheduling for Power-efficient Data Center NetworksabstractPower-efficient Data Center Networks (DCNs) have been proposed to save power of DCNs using OpenFlow. In these DCNs, the OpenFlow controller adaptively turns on and off links and OpenFlow switches to form a minimum-power subnet that satisfies traffic demand. As the subnet changes, flows are scheduled dynamically to routes composed of active switches and links. However, existing flow scheduling schemes could cause undesired results: (1) power inefficiency: due to unbalanced traffic allocation on active routes, extra switches and links may be activated to cater to bursty traffic surges on congested routes, and (2) Quality of Service (QoS) fluctuation: because of the limited flow entry processing ability, switches cannot timely install/delete/update flow entries to properly schedule flows. In this paper, we propose AggreFlow, a dynamic flow scheduling scheme that achieves power efficiency in DCNs and improved QoS using two techniques: Flow-set Routing and Lazy Rerouting. Flow-set Routing achieves load balancing and reduces the number of entry installment on switches by routing flows in a coarse-grained flow-set fashion. Lazy Rerouting maintains load balancing and spreads rerouting operations over a relatively long period of time, reducing the burstiness of entry installment/deletion/update on switches. We built a NS3 based fat-tree network simulation platform to evaluate AggreFlow's performance. The simulation results show AggreFlow reduces power consumption by about 18%, achieves load balancing and improved QoS (i.e., low packet loss rate and reducing the number of processing entries for flow scheduling by 98%), compared with baseline schemes. Zehua Guo 0001, Shufeng Hui, Yang Xu 0010, H. Jonathan Chao |
IWQoS | 1 |
| 2016 | Incremental Switch Deployment for Hybrid Software-Defined NetworksabstractSoftware-Defined Networking (SDN) brings great opportunities to improve network performance. However, due to budget constraints and technique limitations, Internet Service Providers (ISPs) can upgrade only a limited number of conventional switches to SDN switches in real backbone networks at one time. In this paper, we propose one heuristic scheme for deploying SDN switches in hybrid SDNs. Our scheme works for two different cases: (1) maximizing the network control ability with a given upgrading budget constraint, and (2) minimizing the upgrading cost to achieve the best network control ability. We evaluate our scheme in real topologies. We evaluate our scheme in real topologies. The results show that our scheme can achieve 95% of flows controlled with only 10% upgrading cost. Xuya Jia, Yong Jiang 0001, Zehua Guo 0001 |
LCN | 3 |
| 2016 | Reducing and Balancing Flow Table Entries in Software-Defined NetworksabstractSoftware-Defined Networking (SDN) allows flexible and efficient management of networks. However, the limited capacity of flow tables in SDN switches hinders the deployment of SDN. In this paper, we propose a novel routing scheme to improve the efficiency of flow tables in SDNs. To efficiently use the routing scheme, we formulate an optimization problem with the objective to maximize the number of flows in the network, constrained by the limited flow table space in SDN switches. The problem is NP-hard, and we propose the K Similar Greedy Tree (KSGT) algorithm to solve it. We evaluate the performance of KSGT against "traditional" SDN solutions with real-world topologies and traffic. The results show that, compared to the existing solutions, KSGT can reduce about 60% of flow entries when processing the same amount of flows, and improve about 25% of the successful installation and forwarding flows under the same flow table space. Xuya Jia, Yong Jiang 0001, Zehua Guo 0001, Zhenwei Wu |
LCN | 3 |
| 2015 | JumpFlow: Reducing flow table usage in software-defined networks
Zehua Guo 0001, Yang Xu 0010, Marco Cello, Junjie Zhang 0001, Mingjian Liu, H. Jonathan Chao |
Comput. Networks | 1 |
| 2014 | Improving the performance of load balancing in software-defined networks through load variance-based synchronization
Zehua Guo 0001, Mu Su, Yang Xu 0010, Zhemin Duan, Luo Wang, Shufeng Hui, H. Jonathan Chao |
Comput. Networks | 1 |
| 2014 | JET: Electricity cost-aware dynamic workload management in geographically distributed datacenters
Zehua Guo 0001, Zhemin Duan, Yang Xu 0010, H. Jonathan Chao |
Comput. Commun. | 1 |