EDBT 2026 Demo / reviewers in the wild / expert
Songshi Dou
dblp:276/2394
· DBLP profile ↗
34ranked-venue papers
16as first author
33since 2021 · last 2026
0000-0002-6665-0917ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 26 · 10 first-author · 25 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Load-aware Ground Station Assignment for Low Earth Orbit Satellite NetworksabstractLow Earth Orbit (LEO) satellite constellations face ground segment bottlenecks due to uneven user demand, which overloads Ground-to-Satellite Links (GSLs). The common practice of routing traffic to the nearest Ground Station (GS) to minimize latency often causes severe load imbalance. This paper proposes a load-aware assignment strategy that minimizes the maximum GSL utilization by routing traffic to non-nearest GSs via inter-satellite links. To maintain service quality, assignments are constrained by a latency threshold relative to the nearest-GS baseline. We formulate this as a mixed-integer linear program. Preliminary results using realistic Starlink constellation parameters show the proposed solution can reduce average maximum GSL utilization, mitigating ground segment congestion. Songshi Dou, Jinxian Wu, Zehua Guo 0001, Kwan Lawrence Yeung |
CCNC | 1 |
| 2026 | QoS-driven Network Soft Slicing through Token-based Hierarchical Scheduling
Xiaoyang Fu, Zehua Guo 0001, Songshi Dou |
IWQoS | 3 |
| 2026 | Ensuring QoS Stability in Dynamic WANs through Selectively Optimized Range Routing
Zehua Guo 0001, Songshi Dou, Minghao Ye |
IWQoS | 3 |
| 2026 | CEINR: Critical Element-aware In-Network Recovery for Distributed Machine Learning
Jixing Yang, Shuran Zhang, Zehua Guo 0001, Songshi Dou, Jianye Wang |
IWQoS | 4 |
| 2026 | Quest: Quality of Service-Centric Resilient Routing for Software-Defined Wide Area NetworksabstractEmerging network applications pose diverse Quality of Service (QoS) demands, prompting the adoption of Software-Defined Networking (SDN) in Wide Area Networks (WANs), known as Software-Defined Wide Area Networks (SD-WANs). SDN introduces path programmability, allowing controllers to dynamically adjust flow forwarding paths at their switches to accommodate varying traffic conditions and differentiated QoS requirements. However, controller failures can lead to the offline state of controlled switches, causing flows traversing these switches to lose path programmability and degrade QoS. Existing solutions typically focus on indirect metrics such as path programmability or coarse-grained load balancing, which often neglect the individual QoS demands of each flow and the control plane performance. To fill this research gap, we propose Quest, a QoS-aware resilient routing framework designed to preserve QoS in both data and control planes during controller failures. Compared to existing solutions, Quest explicitly addresses the critical limitations by jointly optimizing fine-grained QoS-aware flow routing and control plane latency. We formulate an optimization problem that jointly ensures QoS-aware flow routing and minimizes control plane delays. Due to the complexity of this problem, we utilize a linearization approach and further develop a heuristic QoS-centric resilient routing algorithm. Extensive simulations using real-world topologies and traffic traces demon-strate that Quest substantially improves network performance, achieving a 1.71% increase in average throughput ratio, a 46.65% reduction in average latency, and a 62.74% decrease in average control latency compared to baseline approaches. Songshi Dou, Zehua Guo 0001 |
IEEE Trans. Netw. | 1 |
| 2026 | Maintaining Predictable QoS for Online Service Provisioning in Non-Terrestrial Networks via Safe Transfer LearningabstractEmerging mega-constellations with numerous Low Earth Orbit (LEO) satellites actively provide pervasive Internet services worldwide, which are usually considered crucial components of Non-Terrestrial Networks (NTNs). However, the high mobility and limited coverage of LEO satellites introduce frequent handovers, causing network interruptions and degrading Quality of Service (QoS). While many efforts have been made to alleviate the impact of handovers on service provisioning from NTNs, they usually assume channel conditions are pre-determined and remain unchanged as satellites move, which is different from real situations and thus may experience significant performance degradation compared to theoretical analysis. In this paper, we proposeOracleto promise QoS-aware service provisioning in NTNs under dynamic channel conditions. Specifically, we mathematically formulate a channel model to characterize dynamic channel conditions in NTNs and develop a QoS maximization problem considering handover frequency and transmission capacity. To accommodate the dynamic nature of NTNs, we introduce a Model Predictive Control (MPC)-based controller to predict future network status and generate control strategies correspondingly, and leverage Digital Twin (DT) for real-time network status consideration. For higher efficiency, we further employ Generative Artificial Intelligence (GAI) with a safe transfer learning-based framework to enhance model adaptivity to environmental uncertainties and ensure feasible control decisions in real-world NTNs. Extensive simulation results under the real-world constellation demonstrate thatOraclecan enhance up to$3\times $QoS during online service provisioning. Shengyu Zhang 0003, Songshi Dou, Zhenglong Li 0003, Kwan Lawrence Yeung, Tony Q. S. Quek |
IEEE Trans. Netw. | 2 |
| 2026 | SpaceMeet: Bringing Conferencing Closer Through In-Orbit Conferencing Services
Songshi Dou, Feihu Jin, Jinxian Wu, A-Long Jin, Kwan Lawrence Yeung |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | Oracle: QoS-Aware Online Service Provisioning in Non-Terrestrial Networks with Safe Transfer Learning
Shengyu Zhang 0003, Songshi Dou, Zhenglong Li 0003, Kwan Lawrence Yeung, Tony Q. S. Quek |
INFOCOM | 2 |
| 2025 | Maintaining Predictable Traffic Engineering Performance Under Controller Failures for Software-Defined WANsabstractMany new cloud services and applications have emerged recently. They account for a large share of traffic in Wide Area Networks (WANs) and provide traffic with various Quality of Service (QoS) requirements. Software-Defined Wide Area Network (SD-WAN) offers a promising opportunity for improving the performance of these applications with flexible network management. Nevertheless, SD-WANs are managed by controllers, and unpredictable controller failures may degrade flexible network management. Switches previously controlled by the failed controllers become offline, and flows traversing these offline switches lose the path programmability to route flows on available forwarding paths. Thus, these offline flows cannot be routed/rerouted on available paths to accommodate potential traffic variations, leading to severe performance degradation. Traffic Engineering (TE) is a prevalent network application, which aims to enable differentiable QoS for these numerous cloud services and applications. However, TE performance cannot be guaranteed when controller failures happen due to the loss of flexible network management. Existing recovery solutions reassign offline switches to other active controllers to recover the degraded path programmability but may not promise good TE performance since higher path programmability does not necessarily guarantee satisfactory TE performance. In this paper, we propose ARES to provide predictable TE performance under controller failures. We formulate an optimization problem, which aims to maintain predictable TE performance by jointly considering fine-grained flow-controller reassignment and flow rerouting. Given that the proposed problem is proven to be NP-hard, we further propose a heuristic algorithm to efficiently solve this problem. Specifically, when controller failures occur, ARES updates real-time network information with traffic traces and failure status to calculate optimal flow-controller reassignment and flow rerouting policies. ARES then reassigns and reroutes offline flows to maintain predictable TE performance. Extensive simulation results under two real-world topologies with traffic traces demonstrate that our problem formulation exhibits comparable load balancing performance to optimal TE solution without controller failures, and the proposed ARES can significantly improve average load balancing performance by up to 35.79% with low computation time compared with the state-of-the-art solution. Songshi Dou, Zehua Guo 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2025 | Matchmaker: Maintaining QoS-Aware and Predictable Load Balancing Performance for LEO Mega-Constellations
Songshi Dou, Jinxian Wu, Shengyu Zhang 0003, Xianhao Chen, Tony Q. S. Quek, Kwan Lawrence Yeung |
IEEE Trans. Commun. | 1 |
| 2025 | Toward Improved Performance of Inner Convex Approximation for Suboptimal Nonlinear MPCabstractInner convex approximation is a compelling method that enables the real-time implementation of suboptimal nonlinear model predictive controls (MPCs). However, it suffers from a slow convergence rate, which prevents suboptimal MPC from achieving better performance within a specific sample time. To address this issue, we first reformulate the conventional inner convex approximation procedure as a root-finding problem for a nonlinear equation. Then, under mild assumptions, a comprehensive functional analysis is performed on the derived nonlinear equation, focusing on its continuity, differentiability, and the invertibility of the Jacobian matrix. Building on this analysis, we propose an improved algorithm that applies Broyden's method to accelerate the root-finding procedure of this derived nonlinear equation, thereby enhancing the convergence rate of the conventional inner convex approximation method. We also provide a detailed analysis of the proposed algorithm's convergence properties and computational complexity, showing that it achieves a locally superlinear convergence rate without devoting much additional computational effort. Simulation experiments are performed in an obstacle avoidance scenario, and the results are compared to the conventional inner convex approximation method to assess the effectiveness and advantages of the proposed approach. Jinxian Wu, Li Dai 0001, Songshi Dou, Yunshan Deng, Yuanqing Xia |
IEEE Trans. Cybern. | 3 |
| 2025 | Unleashing the Potential of LEO Constellations in Building Resilient and Low-Latency Control Plane for SD-WANsabstractDelivering seamless network services (e.g., video streaming, AR/VR, and cloud gaming) in Wide Area Networks (WANs) relies on flexible traffic management, which is facilitated by Software-Defined Wide Area Networks (SD-WANs), or SDN in WANs. In SD-WANs, control and data traffic usually share the same links for cost savings, a practice known as in-band control. The SDN controller can periodically send control messages to switches via control channels for routing policy updates to maintain satisfactory network performance. However, when a link failure occurs, control channels become disconnected. Since the controller can no longer communicate with the switches, the desired flexible traffic management cannot be promised. Although backup paths can be preconfigured to reconnect some control channels, high control latency may be introduced due to the meandering routes. Fortunately, commercial Low-Earth Orbit (LEO) mega-constellations, which provide pervasive and low-latency Internet services, present a promising solution to this control resiliency issue. Inspired by the rapid deployment of these LEO constellations, we propose a novel control plane design calledSpaceHelperto leverage the LEO satellite network for improving SD-WANs’ control resiliency.SpaceHelpersmartly integrates the LEO satellite network with terrestrial SD-WAN to reconnect control channels during link failure, which is formulated as an optimization problem to minimize overall control latency. A heuristic algorithm is proposed to solve the problem efficiently and guarantee prompt control channel reconnection. Performance evaluations are conducted using the Starlink constellation and real-world WAN topologies. Compared to the state-of-the-art in-band solution, we show thatSpaceHelpercan not only provide 100% resiliency, but also significantly reduce average control latency by up to 72.6% and 70.2% under GÉANT and Abilene topologies, respectively. Songshi Dou, Zehua Guo 0001, Kwan Lawrence Yeung |
IEEE Trans. Netw. | 1 |
| 2025 | SpaceCache+: Towards Pervasive Content Delivery via Low-Earth Orbit Mega-ConstellationsabstractEmerging Low-Earth Orbit (LEO) mega-constellations face challenges such as limited bandwidth and highly variable user demand, which can degrade network performance and lead to inefficient satellite resource utilization. One promising solution is to enable Content Delivery Networks (CDNs) within LEO satellites by deploying cache-equipped satellites. However, many existing approaches rely on inter-satellite links, which are not widely used in practice and are typically activated only when terrestrial ground station coverage is insufficient. Furthermore, the dynamic coverage patterns of satellites and diverse regional content preferences add to the complexity of efficient CDN deployment in space. To address these challenges, we proposeSpaceCache+, a satellite-based CDN framework. We introduce a new metric,user benefit, that jointly captures user coverage and latency reduction to assess the effectiveness of cache satellite deployment. Recognizing that deployment typically occurs incrementally, we formulate theUser Benefit-centric Cache Satellite Deploymentproblem and design an efficient heuristic solution. To enhance content placement, we also propose a cache replacement policy based on zero-shot meta-learning, which adapts to both regional content popularity and satellite mobility. We evaluate the performance ofSpaceCache+using real-world constellation settings with CDN traces. Compared with benchmark strategies,SpaceCache+improves user benefit and cache hit ratio by up to 66.29% and 77.12%, respectively. Songshi Dou, Shengyu Zhang 0003, Zhenglong Li 0003, Jinxian Wu, Xianhao Chen, Kwan Lawrence Yeung |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | Achieving Predictable and Scalable Load Balancing Performance in LEO Mega-ConstellationsabstractWith the increasing deployment scale of Low Earth Orbit (LEO) mega-constellations, more satellites are expected to become visible to users simultaneously, bringing a new opportunity to optimize the network performance by properly assigning users to satellites. In this paper, we consider LEO mega-constellations without Inter-Satellite Links (ISLs) while assuming there are enough ground relays for inter-satellite communications. To establish a path from a user terminal to the nearest ground station (which serves as a gateway to the Internet), shortest path routing is usually adopted. To focus on the problem of user-satellite assignment, as well as to make routing more scalable, we divide the routing process into two parts: assigning the user terminal to a visible satellite, and finding a path from the satellite to the nearest ground station. For simplicity, shortest path routing is assumed in the second part. Aiming at minimizing the Maximum Satellite Utilization (MSU), a Mixed Integer Linear Programming (MILP), called Optimal User-Satellite Assignment (OUSA), is formulated. Performance evaluations are conducted based on Starlink Phase I mega-constellation and AWS ground station locations. As compared with the existing solutions, we show that the average load balancing performance can be improved by up to 33.29%. Songshi Dou, Shengyu Zhang 0003, Kwan Lawrence Yeung |
ICC | 1 |
| 2024 | Enabling Practical and Pervasive Content Delivery from Emerging LEO Mega-ConstellationsabstractEmerging Low Earth Orbit (LEO) mega-constellations face the challenge of limited bandwidth when providing global Internet services to users. Constructing Content Delivery Networks (CDNs) in LEO mega-constellations is viewed as a feasible solution to address this issue. However, deploying cache satellites, or satellites with CDN servers, for practical and pervasive content delivery is costly and challenging. To overcome this challenge, we consider LEO mega-constellations without inter-satellite links and formulate an integer linear programming problem called User Coverage-aware Cache Satellite Deployment, which aims to maximize the minimum user coverage among all time intervals with a given number of cache satellites. To efficiently solve this problem, a heuristic algorithm called SpaceCache is also proposed. Performance evaluations are conducted based on the Starlink mega-constellation with ground station locations and real-world CDN traces. Compared with benchmark algorithms, we show that the minimum user coverage performance can be improved by up to 41.82% and the cache hit ratio by up to 24.83%. Songshi Dou, Xianhao Chen, Kwan Lawrence Yeung |
ICME | 1 |
| 2024 | ARES: Predictable Traffic Engineering under Controller Failures in SD-WANsabstractEmerging web applications (e.g., video streaming and Web of Things applications) account for a large share of traffic in Wide Area Networks (WANs) and provide traffic with various Quality of Service (QoS) requirements. Software-Defined Wide Area Networks (SD-WANs) offer a promising opportunity to enhance the performance of Traffic Engineering (TE), which aims to enable differentiable QoS for numerous web applications. Nevertheless, SD-WANs are managed by controllers, and unpredictable controller failures may undermine flexible network management. Switches previously controlled by the failed controllers may become offline, and flows traversing these offline switches lose the path programmability to route flows on available forwarding paths. Thus, these offline flows cannot be routed/rerouted on previous paths to accommodate potential traffic variations, leading to severe TE performance degradation. Existing recovery solutions reassign offline switches to other active controllers to recover the degraded path programmability but fail to promise good TE performance since higher path programmability does not necessarily guarantee satisfactory TE performance. In this paper, we propose ARES to provide predictable TE performance under controller failures. We formulate an optimization problem to maintain predictable TE performance by jointly considering fine-grained flow-controller reassignment using P4 Runtime and flow rerouting and propose ARES to efficiently solve this problem. Extensive simulation results demonstrate that our problem formulation exhibits comparable load balancing performance to optimal TE solution without controller failures, and the proposed ARES significantly improves average load balancing performance by up to 43.36% with low computation time compared with existing solutions. Songshi Dou, Zehua Guo 0001 |
WWW | 1 |
| 2024 | Mitigating the impact of controller failures on QoS robustness for software-defined wide area networks
Songshi Dou, Zehua Guo 0001 |
Comput. Networks | 1 |
| 2024 | Byzantine-robust Federated Learning via Cosine Similarity Aggregation
Tengteng Zhu, Zehua Guo 0001, Jiaxin Tan, Songshi Dou, Wenrun Wang, Zhenzhen Han |
Comput. Networks | 5 |
| 2024 | Maintaining the Network Performance of Software-Defined WANs With Efficient Critical RoutingabstractSoftware-Defined Networking (SDN) brings new opportunities to improve network performance of Wide Area Networks (WANs). To enhance the control plane’s processing ability, a Software-Defined Wide Area Network (SD-WAN) usually employs multiple SDN controllers to control a large scale network. The controllers keep a consistent network state with each other via the controller synchronization. A controller synchronization usually involves all controllers, and the network cannot operate until the synchronization is done. Therefore, existing controller synchronization schemes could affect network performance and increase high resource consumption, thus increasing the complexity of running the SD-WAN. In this paper, we propose ThresHold-based critical flOw Routing (THOR) to maintain good load balancing performance of the network. THOR identifies and reroutes critical flows in single domains to achieve the network requirement without affecting other domains. If the rerouting does not satisfy the performance threshold, we generate and reroute several traversing flows among several domains to improve network performance. Simulation results show that THOR achieves 87% of the load balancing performance and reduces the number of synchronizations by approximately 80%, compared with the existing optimal flow routing. Zehua Guo 0001, Songshi Dou, Bida Zhang, Weichao Wu |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2024 | EPIC: Traffic Engineering-Centric Path Programmability Recovery Under Controller Failures in SD-WANsabstractSoftware-Defined Wide Area Networks (SD-WANs) offer a promising opportunity to enhance the performance of Traffic Engineering (TE). With the help of Software-Defined Networking (SDN), TE can promptly respond to traffic changes and maintain network performance by leveraging a global network view. One of the key benefits of SDN for TE is path programmability, which is empowered by SDN controllers to enable dynamic adjustments of flows’ forwarding paths. However, controller failures pose new challenges for SD-WANs since path programmability could be decreased due to the increasing number of offline flows, leading to potential TE performance degradation. Existing recovery solutions mainly focus on recovering path programmability for improving unpredictable network performance but cannot guarantee consistently satisfactory TE performance as expected, since path programmability can only indirectly evaluate network performance. In this paper, we propose EPIC to ensure robust TE performance under controller failures. We observe that frequently rerouted flows could greatly influence TE performance. Enlightened by this, EPIC introduces a novel metric called the TE performance-centric ratio to assess the relevance of different path programmability values for TE performance. The key idea of EPIC lies in identifying frequently rerouted flows during TE operations and prioritizing recovery of the path programmability of these flows under controller failures. We formulate an optimization problem to maximize TE performance-centric path programmability and propose an efficient heuristic algorithm to solve this problem. Evaluation results demonstrate that EPIC can improve average load balancing performance by up to 55.6% compared with baselines. Songshi Dou, Jianye Wang, Zehua Guo 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Toward Improved Path Programmability Recovery for Software-Defined WANs Under Multiple Controller FailuresabstractEnabling path programmability is an essential feature of Software-Defined Networking (SDN). During controller failures in Software-Defined Wide Area Networks (SD-WANs), a resilient design should maintain path programmability for offline flows, which were controlled by the failed controllers. Existing solutions can only partially recover the path programmability rooted in two problems: 1) the implicit preferable recovering flows with long paths and 2) the sub-optimal remapping strategy in the coarse-grained switch level. In this paper, we propose ProgrammabilityGuardian to recover the path programmability of offline flows while maintaining low communication overhead. These goals are achieved through the fine-grained flow-level mappings enabled by existing SDN techniques. ProgrammabilityGuardian configures the flow-controller mappings to recover offline flows with a similar path programmability, maximize the total programmability of the offline flows, and minimize the total communication overhead for controlling these recovered flows. Simulation results of different controller failure scenarios under two different topologies show that ProgrammabilityGuardian recovers offline flows with a balanced path programmability, improves the total programmability of the recovered flows up to 68% and 70%, and reduces the communication overhead by 96% and 99%, compared with the baseline algorithm. Zehua Guo 0001, Songshi Dou, Wenchao Jiang, Yuanqing Xia |
IEEE/ACM Trans. Netw. | 2 |
| 2024 | Maintaining Control Resiliency for Traffic Engineering in SD-WANsabstractSoftware-Defined Networking (SDN) is introduced to Wide Area Networks (WANs) to facilitate network operations and management. One main advantage of SDN is path programmability, i.e., the ability to change forwarding path of flows by controlling underlying SDN switches to accommodate traffic variations. However, the SDN controllers may experience unexpected failure and thus lose its path programmability. The typical solution is to let active controllers control offline switches, which are controlled by failed controllers, by establishing the remapping between the offline switches and active controllers. However, existing remapping solutions do not consider the impact of controller failure on network performance and cannot exhibit predictable network performance under controller failure. In this paper, we take Traffic Engineering (TE) as a typical network scenario and propose Traffic Engineering-Aware Controller-switcH remapping rEcoveRy named TEACHER. We introduce Traffic-aware Path Programmability (TPP) as a new metric to describe the impact of controller failure on TE and design TEACHER based on this metric to smartly recover offline switches. Simulation results show that TEACHER can increase the overall TPP by up to 83.0% and improves the load balancing performance by up to 58.2% under Sprintlink topology with relatively low computation time, compared with baselines. Zehua Guo 0001, Songshi Dou, Jiawei Weng, Xiaoyang Fu, Yuanqing Xia |
IEEE/ACM Trans. Netw. | 3 |
| 2024 | Prophet: Traffic Engineering-Centric Traffic Matrix PredictionabstractTraffic Matrix (TM), which records traffic volumes among network nodes, is important for network operation and management. Due to cost and operation issues, TMs cannot be directly measured and collected in real time. Therefore, many studies work on predicting future TMs based on historical TMs. However, existing works are usually accuracy-centric prediction solutions that mainly focus on improving predicting accuracy of flows’ sizes (i.e., values of elements in TMs) without considering the practical application of TMs. In this paper, we propose a novel TM prediction solution called Prophet for Traffic Engineering (TE), a typical application for TMs which takes TMs as input to optimize routing. We identify that the critical property (i.e., ratio among elements) in a TM plays an important role in TE’s performance. Based on this analysis, we adopt the matrix normalization to maintain the critical property in TMs and customize a TE-centric angle loss function to introduce scale invariance of TMs for capturing the overall relationship error. Different from the element-wise Mean Squared Error (MSE) loss function in accuracy-centric prediction solutions, our proposed TE-centric angle loss function has a clear geometric interpretation, which confines the angle between predicted TM and real TM to zero. Simulation results show that the predicted TMs from Prophet can improve the performance of link-level TE and path-level TE by up to 45.4% and 52.8%, respectively, compared to existing solutions. Yuntian Zhang, Tengteng Zhu, Junjie Zhang 0001, Minghao Ye, Songshi Dou, Zehua Guo 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2023 | RateSheriff: Multipath Flow-aware and Resource Efficient Rate Limiter Placement for Data Center NetworksabstractEmerging cloud services and applications request different Quality of Service (QoS) in Data Center Networks (DCNs). To meet these various requirements, programmable switch-based rate limiters are introduced to provide performance isolation and benefit from easy control and fast deployment. However, existing programmable switch-based rate limiters have two limitations: (1) multipath flows (i.e., MultiPath TCP) cannot be precisely limited, and (2) rate limiter placement solutions in DCNs are missing. These limitations could lead to poor rate limiting performance and low bandwidth utilization. In this paper, we propose RateSheriff to improve rate limiting performance by providing multipath flow-aware and resource efficient rate limiter placement for programmable switch-enabled DCNs. We identify and associate subflows to a multipath flow by extracting and comparing specific packets and header fields. By solving the formulated resource efficient rate limiter placement problem, we can improve rate limiting performance and balance memory utilization among programmable switches in DCNs. Simulation results show that RateSheriff can correctly limit the rate of multipath flows, improve rate limiting performance by up to 46%, and improve memory balancing performance by up to 79% with low computation time, compared with baselines. Songshi Dou, Yongchao He, Sen Liu 0002, Wenfei Wu, Zehua Guo 0001 |
IWQoS | 1 |
| 2023 | Exploring the Impact of Critical Programmability on Controller Placement for Software-Defined Wide Area NetworksabstractControl latency is a critical concern for deploying Software-Defined Networking (SDN) into Wide Area Networks (WANs). A Software-Defined WAN (SD-WAN) can be divided into multiple domains controlled by multiple controllers with a logically centralized view. The control latency is related to the placement of controllers and mappings between switches and controllers. Existing solutions usually consider the propagation delay between switches and controllers as the evaluation metric and fail to consider many important factors of dynamic network states. In this paper, we propose ProgrammabilityExplorer (PE) to optimize the control latency in SD-WAN. Inspired by the selection of critical flows, which have a critical impact on network performance, PE considers the programmability of critical flows at switches and uses this metric to decide the placement of controllers and mappings between switches and controllers. Simulation results show that PE can reduce the control latency by up to 62.3%, 27.5%, 58.3%, and 61.7% under GÉANT, Abilene, Sprintlink, and Tiscali topologies respectively, compared with baseline algorithms. Songshi Dou, Zehua Guo 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | Toward Flexible and Predictable Path Programmability Recovery Under Multiple Controller Failures in Software-Defined WANsabstractSoftware-Defined Networking (SDN) promises good network performance in Wide Area Networks (WANs) with the logically centralized control using physically distributed controllers. In Software-Defined WANs (SD-WANs), maintaining path programmability, which enables flexible path change on flows, is crucial for maintaining network performance under traffic variation. However, when controllers fail, existing solutions are essentially coarse-grained switch-controller mapping solutions and only recover the path programmability of a limited number of offline flows, which traverse offline switches controlled by failed controllers. In this paper, we propose FlexibleProgrammabilityMedic (FlexPM) to provide predictable path programmability recovery under multiple controller failures in SD-WANs. The key idea of FlexPM is to approximately realize flow-controller mappings using hybrid SDN/legacy routing supported by high-end commercial SDN switches. Using the hybrid routing, we can recover programmability by selecting a routing mode for each offline flow at each offline switch in a fine-grained way to fit the given control resource from active controllers and release a few control resource of active controllers by reasonably configuring some normal flows under legacy routing mode. Thus, FlexPM can promise ample control resource to improve the recovery efficiency and further effectively map offline switches to active controllers. Simulation results show that FlexPM outperforms existing switch-level solutions by maintaining balanced programmability and increasing the total programmability of recovered offline flows up to 660% under AT&T topology and 590% under Belnet topology. Zehua Guo 0001, Songshi Dou, Wenfei Wu, Yuanqing Xia |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | Network Coding-based Resilient Routing for Maintaining Data Security and Availability in Software-Defined Networks
Haoran Ni, Zehua Guo 0001, Songshi Dou, Thar Baker |
J. Netw. Comput. Appl. | 4 |
| 2022 | Maintaining Control Resiliency and Flow Programmability in Software-Defined WANs During Controller FailuresabstractProviding resilient network control is a critical concern for deploying Software-Defined Networking (SDN) into Wide-Area Networks (WANs). For performance reasons, a Software-Defined WAN is divided into multiple domains controlled by multiple controllers with a logically centralized view. Under controller failures, we need to remap the control of offline switches from failed controllers to other active controllers. Existing solutions have three limitations: (1) the least flow programmability (e.g., the ability to change paths of flows) cannot be maintained; (2) active controllers could be overloaded, interrupting their normal operations; (3) network performance could be degraded because of the increasing controller-switch communication overhead. In this paper, we propose RetroFlow+ to recover the flow programmability and achieve low communication overhead during controller failures. By intelligently configuring a set of selected offline switches working under the legacy routing mode and several active controllers releasing a few control resources, RetroFlow+ enables active controllers to use the minimum control resource to sustain the flow programmability. RetroFlow+ also smartly transfers the control of offline switches with the SDN routing mode to active controllers to minimize the communication overhead from these offline switches to the active controllers. Simulation results show that RetroFlow+ realizes low communication overhead, recovers all offline flows under one and two controller failures, and improves the flow recovery percentage up to 70% under three controller failures, compared with the state-of-the-art solution. Zehua Guo 0001, Songshi Dou, Sen Liu 0002, Wendi Feng, Wenchao Jiang, Yang Xu 0010, Zhi-Li Zhang |
IEEE/ACM Trans. Netw. | 2 |
| 2021 | ProgrammabilityMedic: Predictable Path Programmability Recovery under Multiple Controller Failures in SD-WANsabstractSoftware-Defined Networking (SDN) promises good network performance in Wide Area Networks (WANs) with the logically centralized control using physically distributed controllers. In Software-Defined WANs (SD-WANs), maintaining path programmability, which enables flexible path change on flows, is crucial for maintaining network performance under traffic variation. However, when controllers fail, existing solutions are essentially coarse-grained switch-controller mapping solutions and only recover the path programmability of a limited number of offline flows, which traverse offline switches controlled by failed controllers. In this paper, we propose ProgrammabilityMedic (PM) to provide predictable path programmability recovery under controller failures in SD-WANs. The key idea of PM is to approximately realize flow-controller mappings using hybrid SDN/legacy routing supported by high-end commercial SDN switches. Using the hybrid routing, we can recover programmability by fine-grainedly selecting a routing mode for each offline flow at each offline switch to fit the given control resource from active controllers. Thus, PM can effectively map offline switches to active controllers to improve recovery efficiency. Simulation results show that PM outperforms existing switch-level solutions by maintaining balanced programmability and increasing the total programmability of recovered offline flows up to 315% under two controller failures and 340% under three controller failures. Songshi Dou, Zehua Guo 0001, Yuanqing Xia |
ICDCS | 1 |
| 2021 | Poster: Enabling Fast Forwarding in Hybrid Software-Defined NetworksabstractEmerging Software-Defined Networking (SDN) technique brings new opportunities to improve network performance. Some SDN-enabled programmable switches are deployed in legacy networks, and thus legacy and programmable switches could coexist, generating hybrid SDNs. In this paper, we study the node upgrade for layer-2 hybrid SDN and propose Shortcutter to accelerate the transmission. Preliminary results show that the proposed Shortcutter can reduce the forwarding path’s length 7% on average, compared with baseline solutions. Yijun Sun, Zehua Guo 0001, Songshi Dou, Junjie Zhang 0001, Xiang Ouyang |
ICNP | 3 |
| 2021 | Video Quality and Popularity-aware Video Caching in Content Delivery NetworksabstractContent Delivery Network (CDN) is a popular service to accelerate object transmission by dynamically caching popular objects at cache points near users. Existing video caching schemes for CDN do not consider some important components of Quality of Experience (QoE). In this paper, we jointly consider video quality and popularity to design a new QoE metric called Video Hit Experience (VHE) and propose an efficient video caching algorithm named Hit ExpeRience-based videO caching (HERO) to improve VHE. Preliminary results show that HERO outperforms existing solutions. Yijun Sun, Zehua Guo 0001, Songshi Dou, Yuanqing Xia |
ICWS | 3 |
| 2021 | Matchmaker: Maintaining network programmability for Software-Defined WANs under multiple controller failures
Songshi Dou, Guochun Miao, Zehua Guo 0001, Weiran Wu, Yuanqing Xia |
Comput. Networks | 1 |
| 2021 | HybridFlow: Achieving Load Balancing in Software-Defined WANs With Scalable RoutingabstractThe scalability issue hinders the deployment of Software-Defined Networking (SDN) in the Wide Area Networks (WANs). Existing solutions have two issues: (1) network performance relies on complicated controller synchronization, which increases the complexity of network control; (2) fine-grained flow processing enables flexible flow control at the cost of high processing load on the controllers and high flow table occupancy on switches. In this paper, we propose a scalable routing solution named HybridFlow, which achieves a good load balancing performance using a single controller with low control overhead (i.e., flow routing and rerouting overhead). HybridFlow mainly employs two techniques: hybrid routing and crucial flow rerouting. Hybrid routing enabled by commercial SDN switches gives us opportunities to reduce the processing load of the controller by routing flows with the hybrid OpenFlow/OSPF mode. Thus, the majority of flows can be routed by OSPF without involving the controller. Crucial flow rerouting realizes load balancing by dynamically identifying crucial flows based on a new metric called Variation Slope and rerouting these flows with the hybrid OpenFlow/OSPF mode. The simulation based on the real traffic traces and network typologies shows that compared with the optimal solution, HybridFlow can achieve 87% of the optimal load balancing performance by rerouting 36% less flows on average. Zehua Guo 0001, Songshi Dou, Yi Wang 0004, Sen Liu 0002, Wendi Feng, Yang Xu 0010 |
IEEE Trans. Commun. | 2 |
| 2020 | Improving the Path Programmability for Software-Defined WANs under Multiple Controller FailuresabstractEnabling path programmability is an essential feature of Software-Defined Networking (SDN). During controller failures in Software-Defined Wide Area Networks (SD-WANs), a resilient design should maintain path programmability for offline flows, which were controlled by the failed controllers. Existing solutions can only partially recover the path programmability rooted in two problems: (1) the implicit preferable recovering flows with long paths and (2) the sub-optimal remapping strategy in the coarse-grained switch level. In this paper, we propose Programmability Guardian to improve the path programmability of offline flows while maintaining low communication overhead. These goals are achieved through the fine-grained flow-level mappings enabled by existing SDN techniques. Programmability Guardian configures the flow-controller mappings to recover offline flows with a similar path programmability, maximize the total programmability of the offline flows, and minimize the total communication overhead for controlling these recovered flows. Simulation results of different controller failure scenarios show that Programmability Guardian recovers all offline flows with a balanced path programmability, improves the total programmability of the recovered flows up to 68%, and reduces the communication overhead up to 83%, compared with the baseline algorithm. Zehua Guo 0001, Songshi Dou, Wenchao Jiang |
IWQoS | 2 |