EDBT 2026 Demo / reviewers in the wild / expert
Shouxi Luo
dblp:151/6230
· DBLP profile ↗
57ranked-venue papers
19as first author
34since 2021 · last 2026
0000-0002-4041-3681ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 36 · 12 first-author · 18 since 2021Systems, architecture and hardware · 12 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Approximate Gradient Synchronization With Adaptive Quantized Gradient Broadcast
Shouxi Luo, Ke Li 0020, Huanlai Xing |
Future Gener. Comput. Syst. | 1 |
| 2026 | Maximizing the computation-communication overlap for distributed deep learning with approximate AllReduce
Shouxi Luo, Gaolin Tang, Huanlai Xing |
Future Gener. Comput. Syst. | 1 |
| 2026 | Maximizing the benefits of in-network aggregation with joint job placement and routing control
Shouxi Luo, Huanlai Xing, Ke Li 0020, Bo Peng 0006 |
Future Gener. Comput. Syst. | 1 |
| 2025 | Efficient In-Network Aggregation With Adaptive Quantization
Zhongxu Su, Shouxi Luo, Ke Li 0020, Huanlai Xing, Bo Peng 0006 |
APNet | 2 |
| 2025 | Dynamic Online Resource Allocation for Synchronization, Retraining, and Inference in Digital Twin NetworkabstractWith the advancement of Intelligent Transportation Systems (ITS), Digital Twin (DT) technology has been widely applied to tasks such as traffic flow modeling and autonomous driving assistance. However, traditional standalone DT models often suffer from weak generalization ability and potential risks of privacy leakage in data-drifting scenarios. To address these challenges, a novel FL-DTN architecture for vehicular networks is proposed by integrating Federated Learning (FL) with Digital Twin Networking (DTN), aiming to preserve data privacy and enhance generalization ability of digital twin models. Specifically, an online resource allocation algorithm, Online Resource Allocation for Synchronization, Retraining and Inference (ORASRI), is designed to dynamically balance the resource allocation among digital twin synchronization, retraining, and inference for Vehicle Digital Twins (VDTs), while adapting to data drift under constrained resource conditions. In addition, a Teacher-Student collaborative mechanism is introduced to improve inference accuracy while reducing resource consumption. Experiments on the MNIST-C dataset show that ORASRI improve 4.6% and 10.3% Inference accuracy in non-FL and FL data-drifting scenarios, respectively. Ke Li 0020, Weichen Tian, Penglin Dai, Shouxi Luo, Huanlai Xing |
GLOBECOM | 5 |
| 2025 | AoI-Error-Aware Data Synchronization for Vehicular Digital TwinabstractSynchronization of vehicular digital twin (VDT) state data is essential for maintaining the accuracy of digital twin models. Existing studies show that VDT data synchronization typically requires a substantial amount of bandwidth and frequent data exchanges. However, real-world vehicle state data are prone to noise interference and bandwidth constraints, causing synchronization errors and significantly degrading VDT accuracy. To evaluate the impact of noise interference on VDT accuracy, we model vehicle state evolution as a discrete-time wiener process and employ a Kalman filter for optimal state estimation. By further analyzing the relationship between estimation error and update timeliness, we find that weighted scheduling based on Age of Information (AoI) effectively suppresses error accumulation and improves synchronization performance. Then, we propose an aoi error-aware scheduling mechanism maximum weighted noise age (MWNA), within a cloud-edge collaborative VDT framework, MWNA dynamically evaluates each vehicle’s state update and prioritizes the transmissions that yield the greatest reduction in estimation error. Compared with baselines, the MWNA policy achieves up to 25–40% lower average error across a range of correlation settings, particularly under high noise correlation and large-scale vehicular scenarios. Ke Li 0020, Xinbang Zhang, Haojun Huang, Shouxi Luo, Huanlai Xing |
GLOBECOM | 5 |
| 2025 | Entangled qubit pricing for quantum networks
Yangming Zhao, Shouxi Luo, Haoze Chen, Chen Tian 0001, Bingheng Yan |
Comput. Networks | 4 |
| 2025 | Knowledge Aggregation Transformer Network for Multivariate Time Series ClassificationabstractOver the years, various sophisticated deep learning algorithms have surfaced for multivariate time series classification (MTSC), notably the dual-network-based model. This model comprises two parallel networks tailored to time series data: one for local feature extraction and the other for global relation extraction. However, effectively integrating these dual networks poses a significant challenge. To address this, we propose a knowledge aggregation transformer network (KATN) for MTSC. KATN, composed of four aggregation transformer blocks, extracts abundant regularizations and connections hidden within the data. Each block incorporates a modified residual network (MResNet) for local feature extraction and a multi-head attention network for global relation extraction. Initially, the block merges MResNet's output feature with that of the multi-head attention network through an additive operation. Subsequently, it aligns features with a fully connected (i.e., dense) layer and activates neural units using the Gaussian error linear unit function. This strategic feature aggregation allows for capturing long-range dependencies among multiple variables in multivariate time series data. Experimental results demonstrate that KATN significantly outperforms 6 state-of-the-art transformer variants, achieving a ‘win’/‘tie’/‘lose’ record of 9/6/15 and securing the lowest AVG_rank score. Furthermore, when evaluated against 18 existing MTSC algorithms across 13 UEA datasets, KATN consistently delivers superior performance, attaining the lowest AVG_rank score among all compared methods. Zhiwen Xiao, Huanlai Xing, Rong Qu, Hui Li 0020, Huagang Tong, Shouxi Luo |
IEEE Trans. Big Data | 6 |
| 2025 | Domain-Specific Transport Protocols for In-Network Processing at the Edge: A Case Study of Accelerating Model SynchronizationabstractNowadays, cross-device federated learning (FL) is the key to achieving personalization services for mobile users and has been widely employed by companies like Google, Microsoft, and Alibaba in production. With the explosive growth in the number of participants, the central FL server, which acts as the manager and aggregator of cross-device model training, would get overloaded, becoming the system bottlenecks. Inspired by the emerging wave of edge computing, an interesting question arises:Could edge clouds help cross-device FL systems overcome the bottleneck?This article provides a cautiously optimistic answer by proposingINP, a FL-specific In-Network Processing framework to achieve the goal. As in-network processing has broken the end-to-end principle of the involved communication and lacks the support of transport protocols, the key is to design domain-specific transport protocols forINP. To fill the gap, we propose the novel Model Download Protocol ofmdpand Model Upload Protocol ofmup. Withmdpandmup, edge cloud nodes along the paths inINPcan easily eliminate duplicated model downloads and pre-aggregate associated gradient uploads for the central FL server, thus alleviating its bottleneck effect, and further accelerating the entire training progress significantly. Shouxi Luo, Pingzhi Fan, Huanlai Xing, Long Luo, Hong-Fang Yu |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | CapMatch: Semi-Supervised Contrastive Transformer Capsule With Feature-Based Knowledge Distillation for Human Activity RecognitionabstractThis article proposes a semi-supervised contrastive capsule transformer method with feature-based knowledge distillation (KD) that simplifies the existing semisupervised learning (SSL) techniques for wearable human activity recognition (HAR), called CapMatch. CapMatch gracefully hybridizes supervised learning and unsupervised learning to extract rich representations from input data. In unsupervised learning, CapMatch leverages the pseudolabeling, contrastive learning (CL), and feature-based KD techniques to construct similarity learning on lower and higher level semantic information extracted from two augmentation versions of the data, "weak" and "timecut," to recognize the relationships among the obtained features of classes in the unlabeled data. CapMatch combines the outputs of the weak- and timecut-augmented models to form pseudolabeling and thus CL. Meanwhile, CapMatch uses the feature-based KD to transfer knowledge from the intermediate layers of the weak-augmented model to those of the timecut-augmented model. To effectively capture both local and global patterns of HAR data, we design a capsule transformer network consisting of four capsule-based transformer blocks and one routing layer. Experimental results show that compared with a number of state-of-the-art semi-supervised and supervised algorithms, the proposed CapMatch achieves decent performance on three commonly used HAR datasets, namely, HAPT, WISDM, and UCI_HAR. With only 10% of data labeled, CapMatch achieves values of higher than 85.00% on these datasets, outperforming 14 semi-supervised algorithms. When the proportion of labeled data reaches 30%, CapMatch obtains values of no lower than 88.00% on the datasets above, which is better than several classical supervised algorithms, e.g., decision tree and -nearest neighbor (KNN). Zhiwen Xiao, Huagang Tong, Rong Qu, Huanlai Xing, Shouxi Luo, Zonghai Zhu, Fuhong Song |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Efficient Parameter Synchronization for Peer-to-Peer Distributed Learning With Selective MulticastabstractRecent advances in distributed machine learning show theoretically and empirically that, for many models, provided that workers will eventually participate in the synchronizations,$i)$the training still converges, even if only$p$workers take part in each round of synchronization, and$ii)$a larger$p$generally leads to a faster rate of convergence. These findings shed light on eliminating the bottleneck effects of parameter synchronization in large-scale data-parallel distributed training and have motivated several optimization designs. In this paper, we focus on optimizing the parameter synchronization forpeer-to-peerdistributed learning, where workers broadcast or multicast their updated parameters to others for synchronization, and proposeSelMcast, a suite of expressive and efficient multicast receiver selection algorithms, to achieve the goal. Compared with the state-of-the-art (SOTA) design, which randomly selects exactly$p$receivers for each worker’s multicast in a bandwidth-agnostic way,SelMcastchooses receivers based on the global view of their available bandwidth and loads, yielding two advantages, i.e., accelerated parameter synchronization for higher utilization of computing resources and enlarged average$p$values for faster convergence. Comprehensive evaluations show thatSelMcastis efficient for both peer-to-peer Bulk Synchronous Parallel (BSP) and Stale Synchronous Parallel (SSP) distributed training, outperforming the SOTA solution significantly. Shouxi Luo, Pingzhi Fan, Ke Li 0020, Huanlai Xing, Long Luo, Hong-Fang Yu |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | ADCC: AoI-aware Decentralized Congestion Control in Cooperative Perception SystemabstractCooperative perception, based on vehicle-to-everything (V2X) communication technology, is a promising solution for connected and automated vehicles (CAVs) to improve their perception capabilities in intelligent transportation systems. The frequency of message transmission in cooperative perception among mobile vehicles plays a crucial role, as it directly impacts communication efficiency, perception accuracy, system response speed, and the safety of real-time applications. Higher transmission frequencies can provide more timely and rich sensory information. This also implies higher consumption of communication resources. However, in a dynamic and complex environment, it is difficult to quantitatively control the message transmission frequency so as to improve the utilization of limited communication resources, ensure the timeliness of perception messages, and maintain fast convergence. To address challenges of timeliness of messages, a timeliness performance metric age of information (AoI) is introduced to control message transmission frequency. This paper deduced the average AoI of the system and designed an AoI-aware decentralized congestion control (ADCC) algorithm for the V2X-based cooperative perception system. Simulation results show that the ADCC algorithm outperforms the classic congestion control algorithm linear adaptive message rate (LIMERIC) in terms of channel utilization and throughput. Specifically, AoI decreased by 47.3%, channel utilization increased by 6.5% and throughput increased by 48.6%. Ke Li 0020, Haojun Huang, Shouxi Luo, Huanlai Xing |
HPCC | 4 |
| 2024 | Towards Optimal Topology-Aware AllReduce SynthesisabstractIn this work, we propose TARS, a Topology-aware AllReduce algorithm Synthesizer, to generate optimal execution plans for AllReduce workloads over arbitrary interconnection network structures. Distinguished from existing topology-aware synthesizers that formulate the two stages of AllReduce (e.g., ReduceScatter-then-AllGather, or Reduce-then-Broadcast) separately, the power of TARS stems from employing a comprehensive Integer Quadratic Programming (IQP) model to formulate the entire workflow precisely. Preliminary studies confirm that, compared with the state-of-the-art scheme, TARS could significantly reduce the completion time of AllReduce. Wenhao Lv, Shouxi Luo, Ke Li 0020, Huanlai Xing |
IWQoS | 2 |
| 2024 | Energy-Efficient Hierarchical Collaborative Learning Over LEO Satellite ConstellationsabstractThe hierarchical collaborative learning within Low Earth Orbit (LEO) satellite constellations, termed LEO-HCL, is gaining increasing popularity by integrating intra-orbit Inter-Satellite Links and orbital edge computing to alleviate the latency issues caused by intermittent satellite connectivity in satellite-ground training architectures. However, LEO-HCL systems are confronted with a triad of challenges: the variable topology induced by satellite mobility, limited onboard computing and communication resources, and stringent energy constraints. In response to these challenges, we propose an energy-efficient training algorithm called FedAAC, which adaptively optimizes both aggregation frequency and model compression ratio within the resource-constrained LEO network. We have conducted a theoretical analysis of model convergence and investigated the relationship between convergence, aggregation frequency, and model compression ratio. Building on this analysis, we offer an approximation algorithm that dynamically calculates the optimal aggregation frequency and compression ratio during the training process. Extensive simulations have demonstrated that FedAAC significantly outperforms existing methods, offering enhanced convergence speed and energy efficiency. Compared to prior solutions, FedAAC achieves a 60% reduction in energy consumption, a 70% decrease in training time, and a 52% lower communication overhead. Long Luo, Chi Zhang 0076, Hong-Fang Yu, Zonghang Li, Gang Sun 0001, Shouxi Luo |
IEEE J. Sel. Areas Commun. | 6 |
| 2024 | Adversarial Reinforcement Learning Based Data Poisoning Attacks Defense for Task-Oriented Multi-User Semantic CommunicationabstractMulti-user semantic communication (MUSC) has emerged as a promising paradigm for future 6G networks and applications, where massive clients (e.g., mobile devices) collaboratively construct a global semantic decoder without sharing their local data. However, due to the lack of direct access to clients’ data, MUSC is vulnerable to data poisoning attacks (DPAs), wherein malicious participants send updates derived from poisoned training samples. Current defense techniques against DPAs are designed for traditional networks and are not directly applicable to MUSC. In this paper, we propose an effective attack-defense game framework, denoted as DPAD-MUSC, tailored to defend against DPAs during image transmission for MUSC. First, we determine each attack-type's optimal attack policy based on reinforcement learning, with the aim of strengthening the attack while avoiding detection. To generate adversarial samples accordingly, we devise an adversarial samples generator (ADV-Generator) based on conditional generative adversarial network (CGAN). Then, we introduce an attack defender (DPA-Defender) to detect data poisoning attacks and exclude poisoned samples from the target model's learning process, with the adversarial samples generated under the guidance of the optimal attack policy to enhance the detector's robustness. Simulation results demonstrate that the DPAD-MUSC can find optimal attack policies that cause a greater accuracy drop in the target model while maintaining a higher evasion rate. The ADV-Generator can generate effective adversarial samples and the DPA-Defender outperforms five state-of-the-art methods on three widely used image datasets under additive white Gaussian noise (AWGN) channel in terms of Top-1 accuracy. Huanlai Xing, Lexi Xu, Shouxi Luo, Penglin Dai, Bowen Zhao 0002, Zhiwen Xiao |
IEEE Trans. Mob. Comput. | 4 |
| 2024 | Releasing the Power of In-Network Aggregation With Aggregator-Aware Routing OptimizationabstractBy offloading partial of the aggregation computation from the logical central parameter servers to network devices like programmable switches, In-Network Aggregation (INA) is a general, effective, and widely used approach to reduce network load thus alleviating the communication bottlenecks suffered by large-scale distributed training. Given the fact that INA would take effects if and only if associated traffic goes through the same in-network aggregator, the key to taking advantage of INA lies in routing control. However, existing proposals fall short in doing so and thus are far from optimal, since they select routes for INA-supported traffic without comprehensively considering the characteristics, limitations, and requirements of the network environment, aggregator hardware, and distributed training jobs. To fill the gap, in this paper, we systematically establish a mathematical model to formulate i) the up-down routing constraints of Clos datacenter networks, ii) the limitations raised by modern programmable switches’ pipeline hardware structure, and iii) the various aggregator-aware routing optimization goals required by distributed training tasks under different parallelism strategies. Based on the model, we develop ARO, an Aggregator-aware Routing Optimization solution for INA-accelerated distributed training applications. To be efficient, ARO involves a suite of search space pruning designs, by using the model’s characteristics, yielding tens of times improvement in the solving time with trivial performance loss. Extensive experiments show that ARO is able to find near-optimal results for large-scale routing optimization in tens of seconds, achieving$1.8\sim 4.0\times $higher throughput than the state-of-the-art solution. Shouxi Luo, Ke Li 0020, Huanlai Xing |
IEEE/ACM Trans. Netw. | 1 |
| 2024 | Efficient Cross-Cloud Partial Reduce With CREWabstractBy allowing$p$out of$n$workers to conductall reduceoperations among them for a round of synchronization,partial reduce, a promising partially-asynchronous variant ofall reduce, has shown its power in alleviating the impacts of stragglers for iterative distributed machine learning (DML). Currentpartial reducesolutions are mainly designed for intra-cluster DML, in which workers are networked with high-bandwidth LAN links. Yet no prior work has looked into the problem of how to achieve efficientpartial reducefor cross-cloud DML, where inter-worker connections are with scarcely-available capacities. To fill the gap, in this paper, we proposeCREW, a flexible and efficient implementation ofpartial reducefor cross-cloud DML. At the high level,CREWis built upon the novel design of employing all active workers along with their internal connection capacities to execute the involved communication and computation tasks; and at the low level,CREWemploys a suite of algorithms to distribute the tasks among workers in a load-balanced way, and deal with possible outages of workers/connections, and bandwidth contention. Detailed performance studies confirm that,CREWnot only shortens the execution of eachpartial reduceoperation, outperforming existing communication schemes such as PS, Ring,TopoAdopt, and BLINK greatly, but also significantly accelerates the training of large models, up to$15\times$and$9\times$, respectively, when compared with the all-to-all direct communication scheme andoriginal partial reducedesign. Shouxi Luo, Renyi Wang, Ke Li 0020, Huanlai Xing |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | Communication-Efficient Federated Learning With Adaptive Aggregation for Heterogeneous Client-Edge-Cloud NetworkabstractClient-edge-cloud Federated Learning (CEC-FL) is emerging as an increasingly popular FL paradigm, alleviating the performance limitations of conventional cloud-centric Federated Learning (FL) by incorporating edge computing. However, improving training efficiency while retaining model convergence is not easy in CEC-FL. Although controlling aggregation frequency exhibits great promise in improving efficiency by reducing communication overhead, existing works still struggle to simultaneously achieve satisfactory training efficiency and model convergence performance in heterogeneous and dynamic environments. This paper proposes FedAda, a communication-efficient CEC-FL training method that aims to enhance training performance while ensuring model convergence through adaptive aggregation frequency adjustment. To this end, we theoretically analyze the model convergence under aggregation frequency control. Based on this analysis of the relationship between model convergence and aggregation frequencies, we propose an approximation algorithm to calculate aggregation frequencies, considering convergence and aligning with heterogeneous and dynamic node capabilities, ultimately achieving superior convergence accuracy and speed. Simulation results validate the effectiveness and efficiency of FedAda, demonstrating up to 4% improvement in test accuracy, 6.8× shorter training time and 3.3× less communication overhead compared to prior solutions. Long Luo, Chi Zhang 0076, Hong-Fang Yu, Gang Sun 0001, Shouxi Luo, Schahram Dustdar |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Densely Knowledge-Aware Network for Multivariate Time Series ClassificationabstractMultivariate time series classification (MTSC) based on deep learning (DL) has attracted increasingly more research attention. The performance of a DL-based MTSC algorithm is heavily dependent on the quality of the learned representations providing semantic information for downstream tasks, e.g., classification. Hence, a model’s representation learning ability is critical for enhancing its performance. This article proposes a densely knowledge-aware network (DKN) for MTSC. The DKN’s feature extractor consists of a residual multihead convolutional network (ResMulti) and a transformer-based network (Trans), called ResMulti-Trans. ResMulti has five residual multihead blocks for capturing the local patterns of data while Trans has three transformer blocks for extracting the global patterns of data. Besides, to enable dense mutual supervision between lower-and higher-level semantic information, this article adapts densely dual self-distillation (DDSD) for mining rich regularizations and relationships hidden in the data. Experimental results show that compared with 5 state-of-the-art self-distillation variants, the proposed DDSD obtains 13/4/13 in terms of “win”/“tie”/“lose” and gains the lowest-AVG_rank score. In particular, compared with pure ResMulti-Trans, DKN results in 20/1/9 regarding win/tie/lose. Last but not least, DKN overweighs 18 existing MTSC algorithms on 10 UEA2018 datasets and achieves the lowest-AVG_rank score. Zhiwen Xiao, Huanlai Xing, Rong Qu, Shouxi Luo, Penglin Dai, Bowen Zhao 0002, Yuan-Shun Dai |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2023 | On ECG Signal Classification: An NAS-empowered Semantic Communication SystemabstractThis paper proposes a task-oriented semantic communication system for electrocardiogram (ECG) signal classification, called ECG-SC-DARTS. Based on deep learning, this system adopts the differentiable neural architecture search (DARTS) to automatically design the neural architecture of the semantic encoder under various channels. This paper improves the performance of the original DARTS by introducing a new recurrent neural network (RNN) cell with residual structure and a noise adding scheme for skip-connections. The RNN cell enhances the temporal semantic information extraction ability while the added noise reduces the risk of performance collapse caused by skip-connections. Experimental results demonstrate that ECGSC-DARTS generates appropriate neural architectures for the semantic encoder under AWGN, Rayleigh and Rician channels and these architectures outperform a number of baseline models, such as the original DARTS, fully convolutional network, multi-layer perception, and ResNet, regarding F1-score. Moreover, ECGSC-DARTS is more reliable than the traditional communication system in harsh channel environment. Huanlai Xing, Huaming Ma, Zhiwen Xiao, Xinhan Wang, Bowen Zhao 0002, Shouxi Luo, Lexi Xu |
TrustCom | 6 |
| 2023 | Evolutionary Multi-Objective Reinforcement Learning Based Trajectory Control and Task Offloading in UAV-Assisted Mobile Edge ComputingabstractThis paper studies the trajectory control and task offloading (TCTO) problem in an unmanned aerial vehicle (UAV)-assisted mobile edge computing system, where a UAV flies along a planned trajectory to collect computation tasks from smart devices (SDs). We consider a scenario that SDs are not directly connected by the base station (BS) and the UAV has two roles to play: MEC server or wireless relay. The UAV makes task offloading decisions online, in which the collected tasks can be executed locally on the UAV or offloaded to the BS for remote processing. The TCTO problem involves multi-objective optimization as its objectives are to minimize the task delay and the UAV's energy consumption, and maximize the number of tasks collected by the UAV, simultaneously. This problem is challenging because the three objectives conflict with each other. The existing reinforcement learning (RL) algorithms, either single-objective RLs or single-policy multi-objective RLs, cannot well address the problem since they cannot output multiple policies for various preferences (i.e. weights) across objectives in a single run. An evolutionary multi-objective RL (EMORL) algorithm is applied to address the TCTO problem. We improve the multi-task multi-objective proximal policy optimization of the original EMORL by retaining all new learning tasks in the offspring population, which can preserve promissing learning tasks. The simulation results demonstrate that the proposed algorithm can obtain more excellent non-dominated policies by striking a balance between the three objectives regarding policy quality, compared with two evolutionary algorithms, two multi-policy RL algorithms, and the original EMORL. Fuhong Song, Huanlai Xing, Xinhan Wang, Shouxi Luo, Penglin Dai, Zhiwen Xiao, Bowen Zhao 0002 |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | Meeting Coflow Deadlines in Data Center Networks With Policy-Based Selective CompletionabstractRecently, the abstraction ofcoflowis introduced to capture the collective data transmission patterns among modern distributed data-parallel applications. During processing, coflows generally act as barriers; accordingly, time-sensitive applications prefer their coflows to complete within deadlines, and deadline-aware coflow scheduling becomes very crucial. Regarding these data-parallel applications, we notice that many of them, includinglarge-scale query systems,distributed iterative training, anderasure codes enabled storage, are able to tolerate loss-bounded incomplete inputs by design. This tolerance indeed brings a flexible design space for the schedule of their coflows: when getting overloaded, the network can trade coflow completeness for the timeliness, and balance the completeness of different coflows on demand. Unfortunately, existing coflow schedulers neglect this tolerance, resulting in inflexible and inefficient bandwidth allocations. In this paper, we explore this fundamental trade-off and design POCO, a POlicy-based COflow scheduler, along with a transport layer enhancement scheme, to achieve customizable selective coflow completion for emerging time-sensitive distributed applications. Internally, POCO employs a suite of novel designs along with admission controls to makeflexible,work-conserving, andperformance-guaranteedrate allocation to online coflow requests very efficiently. Extensive trace-based simulations indicate that POCO is highly flexible and achieves optimal coflow schedules respecting the requirements specified by applications. Shouxi Luo, Pingzhi Fan, Huanlai Xing, Hong-Fang Yu |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | On Jointly Optimizing Partial Offloading and SFC Mapping: A Cooperative Dual-Agent Deep Reinforcement Learning ApproachabstractMulti-access edge computing (MEC) and network function virtualization (NFV) are promising technologies to support emerging IoT applications, especially those computation-intensive. In NFV-enabled MEC environment, service function chain (SFC), i.e., a set of ordered virtual network functions (VNFs), can be mapped on MEC servers. Mobile devices (MDs) can offload computation-intensive applications, which can be represented by SFCs, fully or partially to MEC servers for remote execution. This article studies the partial offloading and SFC mapping joint optimization (POSMJO) problem in an NFV-enabled MEC system, where the data from an incoming task is partitioned into two parts, with one part executed locally and the other offloaded to the edge infrastructure for execution. These two parts are independent of each other, but both need to be processed by the same SFC. The objective is to minimize the average cost in the long term which is a combination of execution delay, MD's energy consumption, and usage charge for edge computing. This problem consists of two closely related decision-making steps, namely task partition and VNF placement, which is highly complex and quite challenging. To address this, we propose a cooperative dual-agent deep reinforcement learning (CDADRL) algorithm, where two agents interact with each other. Simulation results show that the proposed algorithm outperforms three combinations of deep reinforcement learning algorithms with respect to cumulative reward and it overweighs a number of baseline algorithms in terms of execution delay, energy consumption, and usage charge. Xinhan Wang, Huanlai Xing, Fuhong Song, Shouxi Luo, Penglin Dai, Bowen Zhao 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | Approximate Gradient Synchronization with AQGBabstractNo abstract available. Shouxi Luo, Ke Li 0020, Huanlai Xing |
APNet | 2 |
| 2022 | Efficient Partial Reduce Across CloudsabstractNo abstract available. Renyi Wang, Shouxi Luo, Ke Li 0020, Huanlai Xing |
APNet | 2 |
| 2022 | Fast Parameter Synchronization for Distributed Learning with Selective MulticastabstractRecent advances in distributed machine learning show theoretically and empirically that, for many models, provided workers would participate in the synchronizations eventually, i) the training still converges, even if only p workers take part in each round of synchronization, and ii) a larger p generally leads to a faster rate of convergence. These findings shed light on eliminating the bottleneck effects of parameter synchronization in large-scale data-parallel distributed training, having motivated several optimization designs.In this paper, we focus on optimizing the parameter synchronization for peer-to-peer distributed learning, in which workers generally broadcast or multicast their updated parameters to others for synchronization, and propose SELMCAST, an expressive and Pareto-optimal multicast receiver selection algorithm, to achieve the goal. Compared with the state-of-the-art design that randomly selects exactly p receivers for each worker’s multicast in a bandwidth-agnostic way, SELMCAST chooses receivers based on the global view of their available bandwidth and loads, yielding two advantages. Firstly, it could optimize the bottleneck sending rate, thus cutting down the time cost of parameter synchronization. Secondly, when more than p receivers are with sufficient bandwidth, they would be selected as many as possible, bringing benefits to the convergence of training. Extensive evaluations show that SELMCAST is efficient and always achieves near-optimal performance. Shouxi Luo, Pingzhi Fan, Ke Li 0020, Huanlai Xing, Long Luo, Hong-Fang Yu |
ICC | 1 |
| 2022 | Eliminating Communication Bottlenecks in Cross-Device Federated Learning with In-Network Processing at the EdgeabstractNowadays, cross-device federated learning (FL) is the key to achieving personalization services for mobile users and has been widely employed by companies like Google, Microsoft, and Alibaba in production. With the explosive increase of participants, the central FL server, which acts as the manager and aggregator of cross-device model training, would get overloaded, becoming the system bottlenecks. Inspired by the emerging wave of edge computing, an interesting question is: could edge clouds help cross-device FL systems overcome the bottleneck?This article provides a cautiously optimistic answer by proposing INP, an FL-specific In-Network Processing framework, along with the novel Model Download Protocol of MDP and Model Upload Protocol of MUP. With MDP and MUP, edge cloud nodes along the paths in INP can easily eliminate duplicated model downloads and pre-aggregate associated gradient uploads for the central FL server, thus alleviating its bottleneck effect, and further accelerating the entire training progress significantly. Shouxi Luo, Pingzhi Fan, Huanlai Xing, Long Luo, Hong-Fang Yu |
ICC | 1 |
| 2022 | Poster: Selective Reduce for Heterogeneous Distributed TrainingabstractTo improve the performance of partial reduce on synchronizing models for heterogeneous data-parallel distributed training, we explore the idea of selective waiting and worker selection to propose the flexible solution of selective reduce. Our preliminary study shows that with progress- and bandwidth-aware decisions, the proposed partial reduce outperforms the original partial reduce significantly, in terms of both the average synchronization scales and completion times. Shouxi Luo, Ke Li 0020, Huanlai Xing |
ICNP | 2 |
| 2022 | Offloading dependent tasks in multi-access edge computing: A multi-objective reinforcement learning approach
Fuhong Song, Huanlai Xing, Xinhan Wang, Shouxi Luo, Penglin Dai, Ke Li 0020 |
Future Gener. Comput. Syst. | 4 |
| 2022 | SelfMatch: Robust semisupervised time-series classification with self-distillationabstractOver the years, a number of semisupervised deep-learning algorithms have been proposed for time-series classification (TSC). In semisupervised deep learning, from the point of view of representation hierarchy, semantic information extracted from lower levels is the basis of that extracted from higher levels. The authors wonder if high-level semantic information extracted is also helpful for capturing low-level semantic information. This paper studies this problem and proposes a robust semisupervised model with self-distillation (SD) that simplifies existing semisupervised learning (SSL) techniques for TSC, called SelfMatch. SelfMatch hybridizes supervised learning, unsupervised learning, and SD. In unsupervised learning, SelfMatch applies pseudolabeling to feature extraction on labeled data. A weakly augmented sequence is used as a target to guide the prediction of a Timecut-augmented version of the same sequence. SD promotes the knowledge flow from higher to lower levels, guiding the extraction of low-level semantic information. This paper designs a feature extractor for TSC, called ResNet–LSTMaN, responsible for feature and relation extraction. The experimental results show that SelfMatch achieves excellent SSL performance on 35 widely adopted UCR2018 data sets, compared with a number of state-of-the-art semisupervised and supervised algorithms. Huanlai Xing, Zhiwen Xiao, Dawei Zhan, Shouxi Luo, Penglin Dai, Ke Li 0020 |
Int. J. Intell. Syst. | 4 |
| 2022 | Beamer: Stage-Aware Coflow Scheduling to Accelerate Hyper-Parameter Tuning in Deep Learning ClustersabstractTraining a neural network requires retraining the same model many times to search for the configuration of hyper-parameters with the best training result. It is common to launch multiple training jobs and evaluate them in stages. At the completion of each stage, jobs with unpromising configurations will be terminated and jobs with new configurations will start. Each job typically performs distributed training across multiple GPUs, and GPUs periodically synchronize their models over the network. However, model synchronizations of running jobs cause severe network congestion, significantly increasing the stage completion time (SCT) and thus the time to successfully search for the desired configuration. Existing flow schedulers are ineffective to reduce SCT since they are agnostic to training stages. In this paper, we propose a stage-aware coflow scheduling method to minimize the average SCT. In this method, an algorithm is designed to order coflows by considering stage information and then coflows are scheduled according to the order. Mathematical analysis shows that the method achieves the average SCT within 20/3 of the optimal. We implement the method in a real system called Beamer. Extensive testbed experiments and simulations show that Beamer significantly outperforms advanced network designs, such as Sincronia, FIFO-LM, and per-flow fair sharing. Yihong He, Weibo Cai, Pan Zhou 0003, Gang Sun 0001, Shouxi Luo, Hong-Fang Yu, Mohsen Guizani |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2021 | Poster: Accelerate Cross-Device Federated Learning With Semi-Reliable Model Multicast Over The AirabstractTo achieve efficient model multicast for cross-device Federated Learning (FL) over shared wireless channels, we propose SRMP, a transport protocol that performs semi-reliable model multicast over the air by leveraging existing PHY-aided wireless multicast techniques. The preliminary study shows that, with novel designs, SRMP could reduce the communication time involved in each round of training significantly. Yunzhi Lin, Shouxi Luo |
ICNP | 2 |
| 2021 | DGT: A contribution-aware differential gradient transmission mechanism for distributed machine learning
Huaman Zhou, Zonghang Li, Qingqing Cai, Hong-Fang Yu, Shouxi Luo, Long Luo, Gang Sun 0001 |
Future Gener. Comput. Syst. | 5 |
| 2021 | RTFN: A robust temporal feature network for time series classification
Zhiwen Xiao, Xin Xu 0009, Huanlai Xing, Shouxi Luo, Penglin Dai, Dawei Zhan |
Inf. Sci. | 4 |
| 2020 | STDPG: A Spatio-Temporal Deterministic Policy Gradient Agent for Dynamic Routing in SDNabstractDynamic routing in software-defined networking (SDN) can be viewed as a centralized decision-making problem. Most of the existing deep reinforcement learning (DRL) agents can address it, thanks to the deep neural network (DNN) incorporated. However, fully-connected feed-forward neural network (FFNN) is usually adopted, where spatial correlation and temporal variation of traffic flows are ignored. This drawback usually leads to significantly high computational complexity due to large number of training parameters. To overcome this problem, we propose a novel model-free framework for dynamic routing in SDN, which is referred to as spatio-temporal deterministic policy gradient (STDPG) agent. Both the actor and critic networks are based on identical DNN structure, where a combination of convolutional neural network (CNN) and long short-term memory network (LSTM) with temporal attention mechanism, CNN-LSTM-TAM, is devised. By efficiently exploiting spatial and temporal features, CNN-LSTM-TAM helps the STDPG agent learn better from the experience transitions. Furthermore, we employ the prioritized experience replay (PER) method to accelerate the convergence of model training. The experimental results show that STDPG can automatically adapt for current network environment and achieve robust convergence. Compared with a number state-of the-art DRL agents, STDPG achieves better routing solutions in terms of the average end-to-end delay. Zhiwen Xiao, Huanlai Xing, Penglin Dai, Shouxi Luo, Muhammad Azhar Iqbal |
ICC | 5 |
| 2020 | Selective Coflow Completion for Time-sensitive Distributed Applications with PocoabstractRecently, the abstraction of coflow is introduced to capture the collective data transmission patterns among modern distributed data-parallel application. During processing, coflows generally act as barriers; accordingly, time-sensitive applications prefer their coflows to complete within deadlines and deadline-aware coflow scheduling becomes very crucial. Shouxi Luo, Pingzhi Fan, Huanlai Xing, Hong-Fang Yu |
ICPP | 1 |
| 2020 | PSNet: Reconfigurable network topology design for accelerating parameter server architecture based distributed machine learning
Qixuan Jin, Dan Wang 0009, Hong-Fang Yu, Gang Sun 0001, Shouxi Luo |
Future Gener. Comput. Syst. | 6 |
| 2020 | Job scheduling for distributed machine learning in optical WAN
Hong-Fang Yu, Gang Sun 0001, Long Luo, Qixuan Jin, Shouxi Luo |
Future Gener. Comput. Syst. | 6 |
| 2020 | Efficient Multisource Data Delivery in Edge Cloud With Rateless Parallel PushabstractAs the key infrastructure for emerging 5G and Internet-of-Things (IoT) applications, micro data centers would be widely deployed at network edges to provide high-bandwidth low-latency cloud service. In these systems, applications would deliver large-size data objects among servers for various purposes like service deployment, application scale-up, and data duplication on demand. Accordingly, reducing delivery time is crucial for the optimization of service delay and system utilization. To accelerate the delivery, this article proposes a multisource-aware adaptive data transmission solution, Parallel Push (PPUSH), by leveraging the fact that data objects in the cloud are generally replicated among servers by design. At the high level, PPUSH achieves efficient delivery of multisource data by launching multiple push flows in parallel; and at the low level, it decouples transfers from different sources by encoding data objects with rateless RaptorQ code, and further employing novel congestion controls to prioritize the bandwidth allocation of concurrent tasks respecting their remaining sizes. Fluid model analysis along with Mininet-based test and packet-level simulation shows that, unlike DCTCP and other proposals, push is robust to packet loss and achieves provable prioritized bandwidth allocation. Extensive simulation results imply that, with above advantages, PPUSH could achieve very efficient data delivery by making use of all available data sources: for instance, compared with the straightforward design of equal-size task split and fair bandwidth allocation, its adaptive task assignment and prioritized traffic scheduling reduce the average task completion time in a tested scenario by 1.495× and 1.329×, respectively, demonstrating a total improvement of 1.586×, when enabled at the same time. Shouxi Luo, Tie Ma, Pingzhi Fan, Huanlai Xing, Hong-Fang Yu |
IEEE Internet Things J. | 1 |
| 2020 | A Multiobjective Computation Offloading Algorithm for Mobile-Edge ComputingabstractIn mobile-edge computing (MEC), smart mobile devices (SMDs) with limited computation resources and battery lifetime can offload their computing-intensive tasks to MEC servers, thus to enhance the computing capability and reduce the energy consumption of SMDs. Nevertheless, offloading tasks to the edge incurs additional transmission time and thus higher execution delay. This article studies the tradeoff between the completion time of applications and the energy consumption of SMDs in MEC networks. The problem is formulated as a multiobjective computation offloading problem (MCOP), where the task precedence, i.e., ordering of tasks in SMD applications, is introduced as a new constraint in the MCOP. An improved multiobjective evolutionary algorithm based on decomposition (MOEA/D) with two performance enhancing schemes is proposed: 1) the problem-specific population initialization scheme uses a latency-based execution location (EL) initialization method to initialize the EL (i.e., either local SMD or MEC server) for each task and 2) the dynamic voltage and frequency scaling-based energy conservation scheme helps to decrease the energy consumption without increasing the completion time of applications. The simulation results clearly demonstrate that the proposed algorithm outperforms a number of state-of-the-art heuristics and metaheuristics in terms of the convergence and diversity of the obtained nondominated solutions. Fuhong Song, Huanlai Xing, Shouxi Luo, Dawei Zhan, Penglin Dai, Rong Qu |
IEEE Internet Things J. | 3 |
| 2020 | JPAS: Job-progress-aware flow scheduling for deep learning clusters
Pan Zhou 0003, Xinshu He, Shouxi Luo, Hong-Fang Yu, Gang Sun 0001 |
J. Netw. Comput. Appl. | 3 |
| 2020 | Efficient File Dissemination in Data Center Networks With Priority-Based Adaptive MulticastabstractIn today's data center networks (DCN), cloud applications commonly disseminate files from a single source to a group of receivers for service deployment, data replication, software upgrade, etc. For these group communication tasks, recent advantages of software-defined networking (SDN) provide bandwidth-efficient ways-they enable DCN to establish and control a large number of explicit multicast trees on demand. Yet, the benefits of data center multicast are severely limited, since there does not exist a scheme that could prioritize multicast transfers respecting the performance metrics wanted by today's cloud applications, such as pursuing small mean completion times or meeting soft-time deadlines with high probability. To this end, we propose PAM (Priority-based Adaptive Multicast), a preemptive, decentralized, and ready-deployable rate control protocol for data center multicast. At the core, switches in PAM explicitly control the sending rates of concurrent multicast transfers based on their desired priorities and the available link bandwidth. With different policies of priority generation, PAM supports a range of scheduling goals. We not only prototype PAM upon the emerged P4-based programmable switch with novel approximation designs, but also evaluate its performance with ns3-based extensive simulations. Results imply that PAM is ready-deployable; it converges very fast, has negligible impacts on coexisting TCP traffic, and always performs near-optimal priority-based multicast scheduling. Shouxi Luo, Hong-Fang Yu, Ke Li 0020, Huanlai Xing |
IEEE J. Sel. Areas Commun. | 1 |
| 2020 | Online job scheduling for distributed machine learning in optical circuit switch networks
Hong-Fang Yu, Gang Sun 0001, Huaman Zhou, Zonghang Li, Shouxi Luo |
Knowl. Based Syst. | 6 |
| 2019 | Customizable network update planning in SDN
Shouxi Luo, Hong-Fang Yu, Long Luo, Lemin Li |
J. Netw. Comput. Appl. | 1 |
| 2019 | Scalable explicit path control in software-defined networks
Long Luo, Hong-Fang Yu, Shouxi Luo, Zilong Ye, Xiaojiang Du, Mohsen Guizani |
J. Netw. Comput. Appl. | 3 |
| 2017 | Enhancing the reliability of services in NFV with the cost-efficient redundancy schemeabstractNetwork Function Virtualization (NFV) transforms the traditional service provision architecture into a software-based structure. In NFV, Virtual Network Functions Forwarding Graph (VNF FG) [1] is used to describe the logic connections among the VNFs placed on the Physical Network (PN). And end-to-end services are represented as flows traverse the VNF FG. To improve the reliability of the services, redundancy is proposed as an effective method. Existing redundancy method protects the unreliable VNF for end-to-end service independently and embeds the backup without considering the reliability of the PN hardware. Such a scheme ignores the global information of the VNF FG and leads to an inaccurate reliability estimation for the services. These deficiencies result in over backup and the utilization reduction of the underlying resource. To overcome these shortcomings, we propose a novel redundancy scheme. By using the Cost-aware Importance Measure (CIM), the structure of the VNF FG is involved in selecting the qualified backups. Meanwhile, with CIM, the placement decision is made by mapping these backups to the PN nodes with high reliability. In our simulation, our method performs well for cutting down the backup cost for up to 46% with respect to the present algorithms and keeping high cost-efficiency. Weiran Ding, Hong-Fang Yu, Shouxi Luo |
ICC | 3 |
| 2017 | Cotask scheduling in cloud computingabstractComputing frameworks have been widely deployed to support global-scale services. A job typically has multiple sequential stages, where each stage is further divided into multiple parallel tasks. We call the set of all the tasks in a stage of a job a cotask. In this paper, we aim to minimize the average Cotask Completion Time (CCT) in cotask scheduling. To the best of our knowledge, there is no prior work on cotask scheduling for cloud computing. We propose the Cotask Scheduling Scheme (CSS), and take MapReduce as a representative of computing frameworks. CSS schedules cotasks following the Minimum Completion Time First (MCTF) policy, and we prove this problem is NP-hard. We formulate the model using the Integer Linear Programming (ILP), and solve it through an efficient heuristics based on ILP relaxation. Through real trace based simulations, we show that CSS is able to reduce the average CCT by up to 62.20% and 69.93% with traces from our testbed and from a large production cluster respectively. Yangming Zhao, Shouxi Luo, Yi Wang 0021, Sheng Wang 0006 |
ICNP | 2 |
| 2016 | Efficient and Low-Delay Task Scheduling for Big Data Clusters in a Theoretical PerspectiveabstractIn big data clusters, task dispatchers assign arriving tasks to one of many workers (servers) for load balancing. Workers schedule task executions for rapidly completing queueing tasks. Both dispatchers and workers are important for optimizing task/job-completion-time (TCT/JCT). Current dispatchers probe loads on workers before assigning every task/job, which incurs expensive message overheads and significant delays. Besides, they use simple First-In-First-Out (FIFO) scheduling on workers, which further harms their TCT/JCT performance due to head-of-line blocking. In our TASCO scheduler, workers report their loads to dispatchers so that dispatchers avoid to probe them, which significantly reduces expensive overheads and delays on current dispatchers. Motivated by recent observations that more than 60% tasks in big data clusters are recurring with predictable task service time, we also use delay- optimal smallest-task-first (STF) scheduling to improve current simple FIFO scheduling on workers. We also derive the average TCT of TASCO based on its equivalence to an M/G/1/STF queue and the insight that workers reporting loads to dispatchers follows a Poisson process in the large- system limit. Our theories and simulation results demonstrate that the average TCT/JCT of TASCO outperforms state-of-art schedulers from 5.3% to 55.9%. Yuanxiang Gao, Hong-Fang Yu, Shouxi Luo, Shui Yu 0001 |
GLOBECOM | 3 |
| 2016 | Achieving Fast and Lightweight SDN Updates with Segment RoutingabstractIn SDN, forwarding rules are frequently updated to adapt to network dynamics. During the procedure, path consistency needs to be preserved; otherwise, in-flight packets might meet with forwarding errors such as loops and black holes. Despite a large number of suggestions have been proposed, they take either a long duration or have high rule-space overheads, thus fail to be practical for large-scale high dynamic networks. In this paper, we propose FLUS, a Segment Routing (SR) based mechanism, to achieve fast and lightweight path updates. Basically, when a route needs a change, FLUS instantly employs SR to construct its desired new path by concatenating some fragments of the already existing paths. After the actual paths are established, FLUS then shifts incoming packets to them and disables the transitional ones. Such a design helps packets enjoy their new paths immediately without introducing rule-space overheads. This paper presents FLUS's segment allocation, path construction, and the corresponding optimal algorithms in detail. Our evaluation based on real and synthesized networks shows: FLUS can handle up to 92-100% updates using SR in real-time and save 72-88% rule overhead compared to prior methods. Long Luo, Hong-Fang Yu, Shouxi Luo, Mingui Zhang, Shui Yu 0001 |
GLOBECOM | 3 |
| 2016 | Information-agnostic coflow scheduling with optimal demotion thresholdsabstractPrevious coflow scheduling proposals improve the coflow completion time (CCT) over per-flow scheduling based on prior information of coflows, which makes them hard to apply in practice. State-of-art information-agnostic coflow scheduling solution Aalo adopts Discretized Coflow-aware Least-Attained-Service (D-CLAS) to gradually demote coflows from the highest priority class into several lower priority classes when their sent-bytes-count exceeds several predefined demotion thresholds. However, current design standards of these demotion thresholds are crude because they do not analyze the impacts of different demotion thresholds on the average coflow delay. In this paper, we model the D-CLAS system by an M/G/1 queue and formulate the average coflow delay as a function of the demotion thresholds. In addition, we prove the valley-like shape of the function and design the Down-hill searching (DHS) algorithm. The DHS algorithm locates a set of optimal demotion thresholds which minimizes the average coflow delay in the system. Real-data-center-trace driven simulations indicate that DHS improves average CCT up to 6.20× over Aalo. Yuanxiang Gao, Hong-Fang Yu, Shouxi Luo, Shui Yu 0001 |
ICC | 3 |
| 2016 | Decentralized deadline-aware coflow scheduling for datacenter networksabstractThis paper presents D2-CAS, a novel decentralized coflow scheduling system, to minimize the rate of deadline missed coflow for datacenter networks. To design D2-CAS, we first formulate the deadline-missed coflow minimization problem and show its equivalence with the well-known problem of minimizing the late jobs in a concurrent open shop, which is NP-hard in ordinary sense. Inspired by Moore-Hodgso's algorithm (MHA), the optimal solution for minimizing late jobs on a single machine, we design an efficient coflow schedule algorithm, CS-MHA, and further propose its decentralized implementation, D2-CAS. Basically, each sender in D2-CAS periodically runs a part of CS-MHA to get a local suggestion for flow priority assignment, and then multiple senders of a coflow negotiate for an orchestrated priority by leveraging their common data receivers. Via delivering each coflow's packets with its negotiated priority, senders finally carry out efficient deadline-aware coflow scheduling in a decentralized fashion. To the best of our knowledge, this is the first paper that theoretically investigates the deadline-aware coflow scheduling problem, while D2-CAS is the first decentralized solution. Real parameter driven simulations imply that, with the simple yet efficient mechanism, D2-CAS greatly outperforms all existing solutions on reducing the deadline-missed coflows (e.g., outperforms Varys more than 2x). Shouxi Luo, Hong-Fang Yu, Lemin Li |
ICC | 1 |
| 2016 | Towards Practical and Near-Optimal Coflow Scheduling for Data Center NetworksabstractIn current data centers, an application (e.g., MapReduce, Dryad, search platform, etc.) usually generates a group of parallel flows to complete a job. These flows compose a coflow and only completing them all is meaningful to the application. Accordingly, minimizing the average Coflow Completion Time (CCT) becomes a critical objective of flow scheduling. However, achieving this goal in today's Data Center Networks (DCNs) is quite challenging, not only because the schedule problem is theoretically NP-hard, but also because it is tough to perform practical flow scheduling in large-scale DCNs. In this paper, we find that minimizing the average CCT of a set of coflows is equivalent to the well-known problem of minimizing the sum of completion times in a concurrent open shop. As there are abundant existing solutions for concurrent open shop, we open up a variety of techniques for coflow scheduling. Inspired by the best known result, we derive a 2-approximation algorithm for coflow scheduling, and further develop a decentralized coflow scheduling system, D-CAS, which avoids the system problems associated with current centralized proposals while addressing the performance challenges of decentralized suggestions. Trace-driven simulations indicate that D-CAS achieves a performance close to Varys, the state-of-the-art centralized method, and outperforms Baraat, the only existing decentralized method, significantly. Shouxi Luo, Hong-Fang Yu, Yangming Zhao, Sheng Wang 0006, Shui Yu 0001, Lemin Li |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Fast lossless traffic migration for SDN updatesabstractMigration of traffic from one configuration to another is common in SDNs due to node/link failures, network maintenance, policy reconfiguration, intrusion detection, network upgrades, and etc. When the network devices are informed by the controller to execute the traffic migration, it's difficult even impossible to force all network devices to perform the update action in a strict synchronized way. Thus the network is likely to see transient overlapped traffic from both the new configuration and the old one. This kind of overlap may cause overload to those hot spots. This paper reveals the transient congestion problem during traffic migration in an SDN update. According to the observation, it's feasible for the controller to schedule ingress nodes to perform the migration in an order thus the transient congestion is avoided. This scheduling problem is formulated as a Mixed Integer Linear Program (MIP) model. If feasible orders exist for ingress network nodes to perform the migration, the MIP can always find the order that achieves the minimum steps in all possibilities. A heuristic method (named ATOMIP (ATOmic-MIP)) is proposed to speed-up the solving of this MIP. Evaluation based on network topologies observed from real ISPs shows that the lossless migration happens in sub-seconds. Long Luo, Hong-Fang Yu, Shouxi Luo, Mingui Zhang |
ICC | 3 |
| 2015 | Minimizing average coflow completion time with decentralized schedulingabstractIn current data centers, an application (e.g. MapReduce) usually generates a collection of parallel flows sharing a common goal. These flows compose a coflow and only completing them all is meaningful. Accordingly, minimizing the average coflow completion time (CCT) becomes a critical objective for flow scheduling. In this topic, the state-of-the-art centralized method, Varys, achieves a good average CCT; but it has the scalability problem. Alternatively, the only existing decentralized method, Baraat, suffers from the head-of-line blocking problem. To solve these problems, we propose D-CAS, a preemptive, decentralized, coflow-aware scheduling system in this paper. D-CAS pursues coflow-level remaining-time-first (MRTF) principle by leveraging a simple negotiation mechanism between each coflow's data senders and receivers. As the MRTF principle is inherently preemptive and proven to be a near-optimal guideline to minimize average CCT, D-CAS avoids the head-of-line blocking problem and gets good performances. Through extensive simulations, we find that D-CAS achieves a performance close to Varys (gap <; 15%) and outperforms Baraat significantly (about 1.4-4×). Shouxi Luo, Hong-Fang Yu, Yangming Zhao, Bin Wu 0002, Sheng Wang 0006, Lemin Li |
ICC | 1 |
| 2015 | Practical flow table aggregation in SDN
Shouxi Luo, Hong-Fang Yu, Lemin Li |
Comput. Networks | 1 |
| 2014 | Dynamic topology management in optical datacenter networksabstractIn this paper, we study how to manage the topology reconfiguration in OSA-based datacenter networks (DCNs). Though an OSA-based DCN can change its topology to adapt to the traffic matrix and improve the network scalability, it requires too much time (10ms) to reconfigure the topology, which may not only incur a great amount of traffic loss in high throughput low latency DCNs, but also bring much performance degradation to the delay sensitive flows. Therefore, a progressive topology reconfiguration scheme is required to reduce the traffic loss and guarantee the performance of delay sensitive flows. To this end, we first formulate the problem as a mathematical model, and then analyze its feasibility and complexity. Based on these analyses, topology management algorithm (TMA) is proposed to calculate the topology reconfiguration scheme that can maintain the topology connectivity during reconfiguration. By simulation, we find that TMA can reduce the traffic loss during topology reconfiguration by up to 50% in most of the cases and reconfigure topology without traffic loss in some cases. Yangming Zhao, Sheng Wang 0006, Shouxi Luo, Hong-Fang Yu, Shizhong Xu |
GLOBECOM | 3 |
| 2014 | Fast incremental flow table aggregation in SDNabstractIn OpenFlow-based SDN, flow tables are TCAM-hungry and commodity switches suffer from limited concrete flow table size. One method for coping with the limitations is to use aggregation schemes to reduce the number of flow entries required to represent the same forwarding semantics. Unfortunately, the aggregation retards table updates and lengthens the updating time. During which, the data plane is inconsistent with the control plane, forwarding errors such as Reachability Failures, Forwarding Loops, Traffic Isolation and Leakage are prone to occur. Since network updates take place frequently in practice, the aggregation scheme must be efficient enough. In this paper we propose offline FFTA (Fast Flow Table Aggregation) and its online improver iFFTA to shrink the flow table size and to provide practical fast updates. iFFTA is the first online non-prefix aggregation scheme. Extensive experiments demonstrate: (1) FFTA is about 200× faster than the previously published best non-prefix aggregation scheme without loss of compression ratio on offline aggregation; and (2) iFFTA achieves about 3× faster than FFTA on online update incorporations with a loss of an acceptable compression ratio per update. Thus the user could make a combination use of FFTA and iFFTA for table aggregations: call iFFTA usually and recall the efficient FFTA once the switch is running out of concrete flow table space. Shouxi Luo, Hong-Fang Yu, Lemin Li |
ICCCN | 1 |