EDBT 2026 Demo / reviewers in the wild / expert
Long Luo
dblp:56/10016
· DBLP profile ↗
54ranked-venue papers
21as first author
41since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 34 · 15 first-author · 23 since 2021Systems, architecture and hardware · 7 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Speak Your Network: Automated Network Emulation Construction with Cost-Efficient Multi-Agent Orchestration
Daolin Zou, Jie Xu 0004, Long Luo, Hong-Fang Yu |
IWQoS | 3 |
| 2026 | SmartCCL: Learn to Schedule Near-Optimal Collective Communication for GPU Clusters
Long Luo, Jingzhao Xie, Haoxiang Luo, Hong-Fang Yu |
SECON | 2 |
| 2026 | Optimizing Timely Bulk Data Transfers With Hybrid Elastic Cloud ResourcesabstractThe persistent disparity between the slow growth of wide-area network bandwidth and the escalating demand for high-speed bulk data transfer poses a significant challenge for inter-datacenter data transmission. Existing approaches typically rely on optimized algorithms for dedicated networks and often struggle to meet the stringent deadline requirements of modern applications cost-effectively. To overcome these limitations, we propose a novel cloud-accelerated transfer approach that is intrinsically deadline-aware. We present a Profit-driven RelAxation-based TransfEr Scheduling (PRATER) system that leverages cloud proxies and multipath transmission to accelerate bulk data transfer. Specifically, we tackle the problem of profit maximization by jointly optimizing cloud proxy deployment and traffic allocation under transmission deadlines, a problem formally modeled and identified as NP-hard. To efficiently derive near-optimal solutions, we decompose this complex problem into two interconnected subproblems: cost minimization through proxy deployment optimization and revenue maximization via traffic allocation optimization. Our efficient iterative algorithm resolves these subproblems by minimizing cloud proxy and bandwidth costs while maximizing timely data transfer completions, thereby enhancing overall system profitability. Extensive experimental results demonstrate that our approach achieves a significant profit improvement of approximately 2.4×-6.1× compared to state-of-the-art existing algorithms. Long Luo, Yunxiang Zhou, Linjian Yu, Jin Shen, Hong-Fang Yu, Schahram Dustdar |
IEEE Trans. Cloud Comput. | 1 |
| 2025 | Which Cluster Meets My Deadline: a Budget-Aware Scheduler for Distributed Training Jobs in Heterogeneous EnvironmentsabstractTraining deep learning (DL) models demands substantial computational resources, often relying on expensive GPUs in a distributed manner. To meet this demand, cloud providers deploy GPU clusters worldwide to offer users compute instance rental services. These GPU clusters are typically heterogeneous, comprising multiple GPU types with varying computational capabilities, and their prices vary significantly across both GPU instance types and geographic regions. Meanwhile, users often have specific deadlines for training their models. Most existing works focus solely on performance, overlooking price heterogeneity and failing to optimize costs effectively. Given the high cost and time demands of model training, focusing on performance alone is insufficient. In this paper, we aim to balance performance and cost through a DL broker service that maximizes the number of jobs completed within their deadlines under a given budget. We propose CADDS, which jointly optimizes cluster placement and dynamically adjusts GPU type and quantity during training to reduce rental costs. We formulate the scheduling problem as an integer nonlinear programming problem and propose an efficient online approach combining greedy and dynamic programming. Experimental results demonstrate that CADDS significantly outperforms existing approaches, improving the deadline satisfactory ratio within budget limits by$\mathbf{5 9. 4 \%}$to$\mathbf{8 0. 2 \%}$. Long Luo, Zonghang Li, Gang Sun 0001, Hong-Fang Yu |
ICC | 2 |
| 2025 | Poster: LLM Multi-Agent Collaboration for Network Deployment and ManagementabstractThis paper presents the MAPLE framework, which harnesses large language models (LLMs) to facilitate multi-agent collaboration for fully automated deployment and management of large-scale networks. Within MAPLE, a supervisor agent interprets natural language instructions from users, orchestrates specialized agents to execute tasks, and validates outcomes through integration with a network simulation platform. Experimental findings show that MAPLE outperforms single-agent approaches in terms of success rates for topology deployment and service configuration. Moreover, experiments reveal that by adaptively employing LLMs with varying capabilities according to task requirements and inter-agent dependencies, the framework effectively balances task success rates with cost efficiency. Zhengyi Cheng, Chongxi Ma, Mingxuan Tang, Jie Xu 0004, Long Luo, Hong-Fang Yu |
ICNP | 6 |
| 2025 | Poster: Simulation-Guided Strategy Generation for Intent-Aware Distributed LLMs TrainingabstractTraining distributed large language models (LLMs) for diverse user intents is challenging due to a high-dimensional, hybrid strategy space and heterogeneous workloads. We present SG2, a simulation-guided strategy generation framework that tackles intent-aware distributed LLM training by unifying Bayesian Optimization with a simulation-based evaluation loop. By leveraging near-realistic performance and intent-satisfaction metrics in simulation, SG2iteratively refines candidate training strategies before real deployment, significantly reducing exploration cost and failure risk. Experiments on heterogeneous LLM training workloads show that SG2consistently achieves higher intent-satisfaction rates and better multi-objective trade-offs than baselines, demonstrating its effectiveness for intent-aware distributed LLM training. Chongxi Ma, Chengyun Zhang, Long Luo, Hong-Fang Yu |
ICNP | 3 |
| 2025 | Poster: Optimizing Transmission for Privacy-Preserving Edge-Cloud Split LLM InferenceabstractToday’s cloud-centric Large Language Model (LLM) inference can raise privacy concerns, as raw user data is sent over public networks. Split LLM inference offers a promising alternative by handling sensitive stages, such as prefill and input/output decoding, locally at the edge while offloading heavy intermediate decoding layers to the cloud. This approach could enhance privacy, yet it introduces a key challenge: how to transfer large intermediate vectors with ultra-low latency and high reliability over dynamic networks. We explore a potential transmission framework combining (i) dynamic congestion control to adapt flow and reduce queueing/loss, and (ii) adaptive FEC and quantization to balance redundancy with compression. Preliminary results indicate that this approach can achieve lower per-token latency and higher throughput than TCP/QUIC, enabling privacy-preserving LLM services that remain highly responsive and aligned with real-time user experience demands. Yunxiang Zhou, Junzhe Wu, Daolin Zou, Yanan Huang, Long Luo, Hong-Fang Yu |
ICNP | 5 |
| 2025 | Unlocking Agentic AI Service Deployment Complexity: Simulation-Guided Strategy Orchestration and OptimizationabstractDeploying agentic AI services, such as large language models (LLMs) training and inference, presents significant challenges due to their complex, interdependent design across multiple layers of strategy space(framework, system, transport, network). Addressing diverse user intents with limited cross disciplinary expertise further exacerbates this complexity. To overcome these hurdles, we introduce a novel simulation-guided closed-loop (SGCL) strategy orchestration and optimization framework that is inherently intent-aware. Our approach leverages high-fidelity simulators to evaluate candidate deployment strategies, employs a surrogate-based Bayesian optimization engine to guide the closed-loop process, and incorporates a layer-wise caching mechanism to minimize redundant simulations and reduce evaluation overhead. We demonstrate its efficacy in distributed LLM training. Compared to baselines, our method consistently achieves higher intent satisfaction ratio and significantly boosts orchestration efficiency, notably reducing time-to-target by up to 94.4%. These findings highlight SGCL as a practical and effective solution for reliable, cost- and time efficient deployment of agentic AI services. Chongxi Ma, Chengyun Zhang, Long Luo, Weihong Wu, Hong-Fang Yu |
ICPADS | 3 |
| 2025 | Enhancing Mobile Immersive Streaming Experience via Deadline-Aware Scheduling and Learning-Enhanced Congestion ControlabstractMobile immersive streaming applications, such as virtual reality (VR) and augmented reality (AR), impose stringent requirements on network transmission to deliver a seamless user experience. However, existing transmission control methods often fall short in dynamic mobile network environments. Traditional transport-layer protocols typically prioritize network-wide Quality of Service (QoS), overlooking the strict deadlines and varying priorities intrinsic to immersive streaming content, thereby adversely impacting application-specific Quality of Experience (QoE). Furthermore, heuristic-based congestion control methods lack adaptability to rapidly changing network environments, whereas machine learning (ML)-driven approaches, despite their adaptability potential, frequently prove computationally intensive and impractical for deployment on resource-constrained mobile devices. To address these drawbacks, we propose an intelligent transmission control approach that integrates a deadline- and priority-aware scheduler with a learning enhanced congestion control mechanism to optimize immersive streaming content delivery. The scheduling module dynamically prioritizes data packets based on their urgency and importance, timely delivering critical data within strict deadlines as much as possible. Complementing this, the hybrid congestion control module strategically combines lightweight heuristics for baseline performance efficiency with a selectively invoked ML model designed to adaptively adjust the sending rate, efficiently responding to real-time fluctuations in network conditions. Experimental results highlight the effectiveness of our approach, achieving QoE improvements of 15% to 45% compared to existing methods across a range of streaming applications and mobile network scenarios. Long Luo, Yunxiang Zhou, Jin Shen, Haozhe Luo, Weihong Wu, Hong-Fang Yu |
IEEE Internet Things J. | 1 |
| 2025 | Fine-grained forest net primary productivity monitoring: Software system integrating multisource data and smart optimizationabstractAbstract Net primary productivity (NPP) is essential for sustainable resource management and conservation, and it serves as a primary monitoring target in smart forestry systems. The predominant method for NPP inversion involves data collection through terrestrial and satellite sensing systems, followed by parameter estimation using models such as the Carnegie‐Ames‐Stanford Approach (CASA). While this method benefits from low costs and extensive monitoring capabilities, the data derived from multisource sensing systems display varied spatial scale characteristics, and the NPP inversion models cannot detect the impact of data heterogeneity on the outcomes sensitively, reducing the accuracy of fine‐grained NPP inversion. Therefore, this paper proposes a modular system for fine‐grained data processing and NPP inversion. Regarding data processing, a two‐stage spatial‐spectral fusion model based on non‐negative matrix factorization (NMF) is proposed to enhance the spatial resolution of remote sensing data. A spatial interpolation model based on stacking generalization with residual correction is introduced to get raster meteorological data compatible with remote sensing images. Furthermore, we optimize the CASA model with the kernel method to enhance model sensitivity and enrich the spatial details of the inversion results with high resolution. Through validation using real datasets, the proposed fusion and interpolation models have significant advantages over mainstream methods. Furthermore, the correlation coefficient () between the estimated NPP using our improved inversion model and the field‐measured NPP is 0.69, demonstrating the feasibility of this platform in detailed forest NPP monitoring tasks. Weitao Zou, Long Luo, Fangyu Sun, Chao Li 0066, Guangsheng Chen, Weipeng Jing 0001 |
Softw. Pract. Exp. | 2 |
| 2025 | Deadline-Aware Online Job Scheduling for Distributed Training in Heterogeneous ClustersabstractThe explosive growth in training data and model sizes has spurred the adoption of distributed deep learning (DL) in heterogeneous computing clusters. Efficiently scheduling distributed training jobs in such heterogeneous environments while ensuring they meet user-specified deadlines remains a critical challenge. While most existing works focus on reducing job completion time in homogeneous clusters, they pay little attention to meeting job deadlines in heterogeneous clusters. To address this issue, we proposeDancer(Deadline-Aware dyNamiC GPU allocation approach for Efficient Resource utilization), a novel framework that dynamically adjusts not only the number but the type of GPUs assigned to each job throughout its training lifecycle.Danceraims to maximize the number of jobs meeting their deadlines in heterogeneous GPU clusters. It decouples job placement from resource allocation and formulates the scheduling optimization problem for maximizing the number of deadline-meeting jobs as an Integer Linear Programming (ILP) problem. To solve this ILP problem in real-time, we propose an online algorithm with a competitive ratio guarantee, leveraging primal-dual and dynamic programming techniques. Extensive trace-driven simulations based on real-world DL workloads demonstrate thatDancersignificantly outperforms state-of-the-art approaches, improving the deadline satisfactory ratio up to 58.9%–74.2%. Long Luo, Gang Sun 0001, Hong-Fang Yu, Bo Li 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2025 | Domain-Specific Transport Protocols for In-Network Processing at the Edge: A Case Study of Accelerating Model SynchronizationabstractNowadays, cross-device federated learning (FL) is the key to achieving personalization services for mobile users and has been widely employed by companies like Google, Microsoft, and Alibaba in production. With the explosive growth in the number of participants, the central FL server, which acts as the manager and aggregator of cross-device model training, would get overloaded, becoming the system bottlenecks. Inspired by the emerging wave of edge computing, an interesting question arises:Could edge clouds help cross-device FL systems overcome the bottleneck?This article provides a cautiously optimistic answer by proposingINP, a FL-specific In-Network Processing framework to achieve the goal. As in-network processing has broken the end-to-end principle of the involved communication and lacks the support of transport protocols, the key is to design domain-specific transport protocols forINP. To fill the gap, we propose the novel Model Download Protocol ofmdpand Model Upload Protocol ofmup. Withmdpandmup, edge cloud nodes along the paths inINPcan easily eliminate duplicated model downloads and pre-aggregate associated gradient uploads for the central FL server, thus alleviating its bottleneck effect, and further accelerating the entire training progress significantly. Shouxi Luo, Pingzhi Fan, Huanlai Xing, Long Luo, Hong-Fang Yu |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | Efficient Parameter Synchronization for Peer-to-Peer Distributed Learning With Selective MulticastabstractRecent advances in distributed machine learning show theoretically and empirically that, for many models, provided that workers will eventually participate in the synchronizations,$i)$the training still converges, even if only$p$workers take part in each round of synchronization, and$ii)$a larger$p$generally leads to a faster rate of convergence. These findings shed light on eliminating the bottleneck effects of parameter synchronization in large-scale data-parallel distributed training and have motivated several optimization designs. In this paper, we focus on optimizing the parameter synchronization forpeer-to-peerdistributed learning, where workers broadcast or multicast their updated parameters to others for synchronization, and proposeSelMcast, a suite of expressive and efficient multicast receiver selection algorithms, to achieve the goal. Compared with the state-of-the-art (SOTA) design, which randomly selects exactly$p$receivers for each worker’s multicast in a bandwidth-agnostic way,SelMcastchooses receivers based on the global view of their available bandwidth and loads, yielding two advantages, i.e., accelerated parameter synchronization for higher utilization of computing resources and enlarged average$p$values for faster convergence. Comprehensive evaluations show thatSelMcastis efficient for both peer-to-peer Bulk Synchronous Parallel (BSP) and Stale Synchronous Parallel (SSP) distributed training, outperforming the SOTA solution significantly. Shouxi Luo, Pingzhi Fan, Ke Li 0020, Huanlai Xing, Long Luo, Hong-Fang Yu |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Analysis and Optimization for Passive One-way Delay Measurement Tax in Container NetworksabstractContainer networks have become crucial to the overall performance and health of network systems due to the increasing adoption of container technologies. Passive per-packet and per-hop one-way delay (p4 h-OWD) measurement is essential for prompt and accurate detection and localization of network issues in container environments. Existing measurement methods for generic virtual networks require intrusive packet modification to uniquely match packets during p4 h-OWD measurements. However, the overhead and ensuing performance impact induced by high-frequency packet operations during the measurement process have been largely overlooked in the context of lightweight and weak-isolation container networks. This paper demonstrates that even state-of-the-art technologies leveraging the efficient extended Berkeley Packet Filter (eBPF) can introduce significant overhead, impacting both the networking and computing performance of container networks during p4 h-OWD measurements. To effectively reduce the p4 h-OWD measurement tax, which encompasses the measurement overhead and its consequent impact, we propose a non-intrusive method to obtain unique packet identifiers in container networks, thereby avoiding the significant operational overhead associated with intrusive packet matching. Building on this foundation, we present CNDMeas, an efficient eBPF -based technology designed for p4 h-OWD mea-surement within container networks. Evaluation results show that CNDMeas effectively limits the increase in CPU time dedicated to handling software interrupts to within 3 % and reduces the impact on both networking performance by up to 71 % in terms of delay, and computing performance during the measurement process compared to state-of-the-art technologies. Jingzhao Xie, Chongxi Ma, Hong-Fang Yu, Long Luo, Gang Sun 0001 |
CLOUD | 4 |
| 2024 | Approximation Algorithms for Minimizing Congestion in Demand-Aware NetworksabstractEmerging reconfigurable optical communication technologies allow to enhance datacenter topologies with demand-aware links optimized towards traffic patterns. This paper studies the algorithmic problem of jointly optimizing topology and routing in such demand-aware networks to minimize congestion, along two dimensions: (1) splittable or unsplittable flows, and (2) whether routing is segregated, i.e., whether routes can or cannot combine both demand-aware and demand-oblivious (static) links.For splittable and segregated routing, we show that the problem is generally 2-approximable, but APX-hard even for uniform demands induced by a bipartite demand graph. For unsplittable and segregated routing, we establish upper and lower bounds of O (log m/ log log m) and Ω (log m/ log log m), respectively, for polynomial-time approximation algorithms, where m is the number of static links. We further reveal that under un-/splittable and non-segregated routing, even for demands of a single source (resp., d estina tion), the problem cannot be approximated better than $\Omega \left({\frac{{{c_{\max }}}}{{{c_{\min }}}}}\right)$ unless P=NP, where cmax(resp., cmin) denotes the maximum (resp., minimum) capacity. It remains NP-hard for uniform capacities, but is tractable for a single commodity and uniform capacities.Our trace-driven simulations show a significant reduction in network congestion compared to existing solutions. Wenkai Dai, Michael Dinitz, Klaus-Tycho Förster, Long Luo, Stefan Schmid 0001 |
INFOCOM | 4 |
| 2024 | Klonet: an Easy-to-Use and Scalable Platform for Computer Networks Education
Tie Ma, Long Luo, Hong-Fang Yu, Xi Chen 0026, Jingzhao Xie, Chongxi Ma, Yunhan Xie, Gang Sun 0001, Tianxi Wei, Li Chen 0008, Yanwei Xu 0004, Nicholas Zhang |
NSDI | 2 |
| 2024 | Energy-Efficient Hierarchical Collaborative Learning Over LEO Satellite ConstellationsabstractThe hierarchical collaborative learning within Low Earth Orbit (LEO) satellite constellations, termed LEO-HCL, is gaining increasing popularity by integrating intra-orbit Inter-Satellite Links and orbital edge computing to alleviate the latency issues caused by intermittent satellite connectivity in satellite-ground training architectures. However, LEO-HCL systems are confronted with a triad of challenges: the variable topology induced by satellite mobility, limited onboard computing and communication resources, and stringent energy constraints. In response to these challenges, we propose an energy-efficient training algorithm called FedAAC, which adaptively optimizes both aggregation frequency and model compression ratio within the resource-constrained LEO network. We have conducted a theoretical analysis of model convergence and investigated the relationship between convergence, aggregation frequency, and model compression ratio. Building on this analysis, we offer an approximation algorithm that dynamically calculates the optimal aggregation frequency and compression ratio during the training process. Extensive simulations have demonstrated that FedAAC significantly outperforms existing methods, offering enhanced convergence speed and energy efficiency. Compared to prior solutions, FedAAC achieves a 60% reduction in energy consumption, a 70% decrease in training time, and a 52% lower communication overhead. Long Luo, Chi Zhang 0076, Hong-Fang Yu, Zonghang Li, Gang Sun 0001, Shouxi Luo |
IEEE J. Sel. Areas Commun. | 1 |
| 2024 | Short-Term Probabilistic Load Forecasting Using Quantile Regression Neural Network With Accumulated Hidden Layer Connection StructureabstractThe integration of distributed energy systems into grids increases the uncertainty of electric loads. Accurate short-term load forecasting is critical to cope with the uncertainty and secure the operation of power systems. In this article, we propose a short-term probabilistic load forecasting model based on the quantile regression neural network (QRNN) and an accumulated hidden layer connection (AHLC) structure. The AHLC structure connects the hidden layers of all the predicted hours and can provide more information to the model output layers. This AHLC structure, together with parallel prediction structure and 1-D convolutional structure, improves the accuracy of the short-term probabilistic load prediction. Adaptive fuzzy control is employed to rectify data anomalies caused by emergency situations. The proposed model has been evaluated using the publicly available GEFCom2014 dataset, the ISO-NE dataset, and the Malaysia dataset. Numerical results show that the proposed AHLC-QRNN model has better performance compared to existing models. Long Luo, Jizhe Dong, Weizhe Kong, Qi Zhang 0085 |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | Accelerating Geo-Distributed Machine Learning With Network-Aware Adaptive Tree and Auxiliary RouteabstractDistributed machine learning is becoming increasingly popular for geo-distributed data analytics, facilitating the collaborative analysis of data scattered across data centers in different regions. This paradigm eliminates the need for centralizing sensitive raw data in one location but faces the significant challenge of high parameter synchronization delays, which stems from the constraints of bandwidth-limited, heterogeneous, and fluctuating wide-area networks. Prior research has focused on optimizing the synchronization topology, evolving from starlike to tree-based structures. However, these solutions typically depend on regular tree structures and lack an adequate topology metric, resulting in limited improvements. This paper proposes NetStorm, an adaptive and highly efficient communication scheduler designed to speed up parameter synchronization across geo-distributed data centers. First, it establishes an effective metric for optimizing a multi-root FAPT synchronization topology. Second, a network awareness module is developed to acquire network knowledge, aiding in topology decisions. Third, a multipath auxiliary transmission mechanism is introduced to enhance network awareness and facilitate multipath transmissions. Lastly, we design policy consistency protocols to guarantee seamless updates of transmission policies. Empirical results demonstrate that NetStorm significantly outperforms distributed training systems like MXNET, MLNET, and TSEngine, with a speedup of 6.5~9.2 times over MXNET. Zonghang Li, Wenjiao Feng, Weibo Cai, Hong-Fang Yu, Long Luo, Gang Sun 0001, Hongyang Du 0001, Dusit Niyato |
IEEE/ACM Trans. Netw. | 5 |
| 2024 | Communication-Efficient Federated Learning With Adaptive Aggregation for Heterogeneous Client-Edge-Cloud NetworkabstractClient-edge-cloud Federated Learning (CEC-FL) is emerging as an increasingly popular FL paradigm, alleviating the performance limitations of conventional cloud-centric Federated Learning (FL) by incorporating edge computing. However, improving training efficiency while retaining model convergence is not easy in CEC-FL. Although controlling aggregation frequency exhibits great promise in improving efficiency by reducing communication overhead, existing works still struggle to simultaneously achieve satisfactory training efficiency and model convergence performance in heterogeneous and dynamic environments. This paper proposes FedAda, a communication-efficient CEC-FL training method that aims to enhance training performance while ensuring model convergence through adaptive aggregation frequency adjustment. To this end, we theoretically analyze the model convergence under aggregation frequency control. Based on this analysis of the relationship between model convergence and aggregation frequencies, we propose an approximation algorithm to calculate aggregation frequencies, considering convergence and aligning with heterogeneous and dynamic node capabilities, ultimately achieving superior convergence accuracy and speed. Simulation results validate the effectiveness and efficiency of FedAda, demonstrating up to 4% improvement in test accuracy, 6.8× shorter training time and 3.3× less communication overhead compared to prior solutions. Long Luo, Chi Zhang 0076, Hong-Fang Yu, Gang Sun 0001, Shouxi Luo, Schahram Dustdar |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | Maximizing Aggregation Throughput for Distributed Training with Constrained In-Network ComputingabstractDistributed training (DT) has become an important and popular practice for collaborative training of high-quality machine learning (ML) models. The communication efficiency of gradient aggregation has been shown to be the primary performance bottleneck for distributed training today. Advanced programmable switches with in-network computing capabilities provide a promising direction for improving the communication efficiency of DT by offloading some gradient aggregations from the host to switches in the network. In this paper, we propose SPAR to optimize the performance of gradient aggregation under constrained in-network computing capabilities. To improve the aggregation throughput, SPAR jointly optimizes the deployment of in-network aggregation switches and the routing of aggregation requests from workers. We formulate this joint optimization problem as an integer nonlinear programming problem and design an efficient greedy algorithm to compute solutions quickly. The experimental results show that SPAR significantly outper-forms the other state-of-the-art solutions based on in-network aggregation, improving aggregation throughput by up to 3×. Long Luo, Shulin Yang, Hong-Fang Yu, Bo Lei 0002 |
ICC | 1 |
| 2023 | High Throughput Routing Path Selection for Payment Channel NetworkabstractThe Payment Channel Network (PCN) has emerged as a prominent solution for addressing the scalability limitations of cryptocurrencies. A critical aspect of PCN transactions is the selection of appropriate routing paths between parties. However, the success rate of payment routing in PCNs is significantly influenced by channel capacity and other payments, making path selection in PCNs more challenging than in traditional networks. In this paper, we present a novel path selection algorithm for PCNs that jointly considers path delay, channel capacity, and inter-path impact to achieve high throughput. Our simulation results demonstrate that the proposed algorithm improves throughput by 24.2% compared to the traditional widest path algorithm, and by 5.3% compared to the traditional shortest path algorithm. Qingqing Cai, Gang Sun 0001, Hong-Fang Yu, Long Luo |
ISCC | 4 |
| 2023 | FedGSync: Jointly Optimized Weak Synchronization and Gradient Transmission for Fast Distributed Machine Learning in Heterogeneous WANabstractDue to privacy and cost reasons, distributed machine learning in Wide-Area Networks(DML-WAN) is becoming an emerging and popular collaborative learning paradigm. However, heterogeneity in computing power and data distribution among workers in different locations has a dramatic impact on training performance, including convergence speed and learning accuracy. Most of the existing works on distributed training mechanisms either focus on computing heterogeneity or data heterogeneity, and none of them can handle both well. In this paper, we propose FedGSync, a novel distributed training mechanism to improve the training performance for DML-WAN, where computing heterogeneity and data heterogeneity usually coexist. To speed up training and improve model accuracy, FedGSync clusters workers into groups according to the similarity of their data distribution and introduce group-based weak synchronization to minimize the synchronization delays waiting for slow workers and the accuracy loss by balancing the contributions of all data distributions. To preserve data privacy and improve efficiency, FedGSync only groups workers based on principal components of gradients and design an approximate grouping mechanism based on Kmeans. To further reduce synchronization time, FedGSync prioritizes packets and uses differential transmission for gradient packets between groups. Evaluation results demonstrate that FedGSync improves convergence speed and learning accuracy under the coexistence of computing heterogeneity and data heterogeneity compared with state-of-the-art distributed training mechanisms. Huaman Zhou, Yihong He, Long Luo, Hong-Fang Yu, Gang Sun 0001 |
SMC | 4 |
| 2023 | HFedMS: Heterogeneous Federated Learning With Memorable Data Semantics in Industrial MetaverseabstractFederated Learning (FL), as a rapidly evolving privacy-preserving collaborative machine learning paradigm, is a promising approach to enable edge intelligence in the emerging Industrial Metaverse. Even though many successful use cases have proved the feasibility of FL in theory, in the industrial practice of Metaverse, the problems of non-independent and identically distributed (non-i.i.d.) data, learning forgetting caused by streaming industrial data, and scarce communication bandwidth remain key barriers to realize practical FL. Facing the above three challenges simultaneously, this paper presents a high-performance and efficient system namedHFedMSfor incorporating practical FL into Industrial Metaverse.HFedMSreduces data heterogeneity through dynamic grouping and training mode conversion (Dynamic Sequential-to-Parallel Training, STP). Then, it compensates for the forgotten knowledge by fusing compressed historical data semantics and calibrates classifier parameters (Semantic Compression and Compensation, SCC). Finally, the network parameters of the feature extractor and classifier are synchronized in different frequencies (Layer-wise Alternative Synchronization Protocol, LASP) to reduce communication costs. These techniques make FL more adaptable to the heterogeneous streaming data continuously generated by industrial equipment, and are also more efficient in communication than traditional methods (e.g., Federated Averaging). Extensive experiments have been conducted on the streamed non-i.i.d. FEMNIST dataset using 368 simulated devices. Numerical results show thatHFedMSimproves the classification accuracy by at least 6.4% compared with 8 benchmarks and saves both the overall runtime and transfer bytes by up to 98%, proving its superiority in precision and efficiency. Shenglai Zeng, Zonghang Li, Hong-Fang Yu, Long Luo, Bo Li 0001, Dusit Niyato |
IEEE Trans. Cloud Comput. | 5 |
| 2023 | Cost-Efficient Scheduling of Multicast Transfers With Deadline Guarantees Across Edge DatacentersabstractModerate-scale datacenters are increasingly deployed at the network edge to support low-latency and high-bandwidth internet of Things (IoT) and 5G applications. These applications usually have bulk data to transfer from one datacenter to several datacenters, which can cause a significant amount of bandwidth costs. Traffic engineering (TE) systems at the network edge must minimize the bandwidth costs of multicast transfers across edge datacenters. To avoid service quality degradation, TE should also guarantee transfer deadlines, which have received little attention in existing work. Due to the dynamic bandwidth pricing schemes, it is challenging to strike a good balance between minimizing costs and guaranteeing transfer deadlines. In this paper, we present CDScheduler a cost-efficient scheduling solution for multicast transfers with deadline guarantees. To reduce bandwidth costs, CDScheduler uses Steiner trees for forwarding and adaptive routing that considers bandwidth price variation and transfer demands. We formulate the cost-efficient multicast transfer scheduling problem and propose an algorithm based on linear program relaxation and randomized rounding to find a solution that guarantees timely completion and reduces cost. Extensive evaluations show that CDScheduler can reduce bandwidth cost significantly and outperforms the state-of-the-art solutions by cutting down up to 78% bandwidth cost. Long Luo, Qixuan Jin, Jingzhao Xie, Gang Sun 0001, Hong-Fang Yu |
IEEE Trans. Serv. Comput. | 1 |
| 2023 | NBSync: Parallelism of Local Computing and Global Synchronization for Fast Distributed Machine Learning in WANsabstractRecently, due to privacy concerns, distributed machine learning in Wide-Area Networks (DML-WANs) attracts increasing attention and has been widely deployed to promote the widespread application of intelligence services that rely on geographically distributed data. DML-WANs is essentially performing collaboratively federated learning over a combination of servers at both edge and cloud on a large spatial scale. However, efficient model training is challenging for DML-WANs because it is blocked by the high overhead of model parameter synchronization between computing servers over WANs. The reason is that there has a sequential dependency between local model computing and global model synchronization of traditional DML-WANs training methods intrinsically producing a sequential blockage between them, e.g., FedAvg. When the computing heterogeneity and the low WAN bandwidth coexist, a long block of global model synchronization prolongs the training time and leads to low utilization of local computing. Despite many efforts on alleviating synchronization overhead with novel communication technologies and synchronization methods, they still use traditional training patterns with sequential dependency and thereby have very limited improvements, such as FedAsync and ESync. In this article, we propose NBSync, a novel training algorithm for DML-WANs, which greatly speeds up the model training by the parallelism of local computing and global synchronization. NBSync employs a well-designed pipelining scheme, which can properly relax the sequential dependency of local computing and global synchronization and process them in parallel so as to overlap their operating overhead in the time dimension. NBSync also realizes flexible, differentiated and dynamical local computing for workers to maximize the overlap ratio in dynamically heterogeneous training environments. Convergence analysis shows that the convergence rate of NBSync training process is asymptotically equal to that of SSGD, and NBSync has a better convergence efficiency. We implemented the prototype of NBSync based on a popular parameter server system, i.e., MXNET's PS-LITE library, and evaluate its performance on a DML-WANs testbed. Experimental results show that NBSync speeds up training about 1.43×–2.79× than state-of-the-art distributed training algorithms (DTAs) in DML-WANs scenarios where computing heterogeneity and low WAN bandwidth coexist. Huaman Zhou, Zonghang Li, Hong-Fang Yu, Long Luo, Gang Sun 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2022 | MMOS: Multi-Staged Mutation Operator Scheduling for Deep Learning Library TestingabstractThe rapid development of deep learning (DL) technology has made deep learning libraries such as Tensor Flow widely used in practice. However, the complexity of DL libraries inevitably leads to multiple vulnerabilities. Recently, research on DL library testing generally designs a variety of mutation operators to generate new models from the perspective of model mutation. However, we find these research efforts do not consider that different mutation operators have different vulnerability mining efficiencies, thus treating all mutation operators equally can lead to inefficiency. Based on the above observation, we design a novel mutation operator scheduling strategy to improve the efficiency of vulnerability mining in the DL library, including the multi-staged mutation operator selection strategy and mutation operator energy allocation strategy. To evaluate the efficiency of our work, we implement a prototype called MMOS. The results show that MMOS finds six more crash bugs, two more NaN bugs and one more inconsistency bug in four widely used DL libraries including TensorFlow, Theano, CNTK and MXNet compared with the random strategy. Among them, MMOS finds 5 previously unknown vulnerabilities in MXNet and 1 in CNTK. Moreover, MMOS outperforms LEMON in vulnerability mining of DL libraries. In total, MMOS finds 13 more vulnerabilities than LEMON. Senyi Li, Long Luo, Hong-Fang Yu |
GLOBECOM | 4 |
| 2022 | vNetRadar: Lightweight and Network-Wide Traffic Measurement in Virtual NetworksabstractMeasuring traffic metrics is indispensable in virtual networks as it is the basis for a wide range of applications, such as network diagnostics and performance evaluation of the network algorithms. However, existing measurement schemes fail to have all these excellent characteristics simultaneously: 1) fine-grained, i.e. to obtain per packet level information. 2) lightweight, namely low CPU and bandwidth overhead. 3) network-wide, which means obtaining metrics of the whole network, e.g. per packet path. 4) easy-to-deploy, which refers to deployment without additional modification of Maximum Transmission Units (MTUs). We design vNetRadar, a virtual network measurement system, which has these excellent characteristics simultaneously. Specifically, vNetRadar 1) identifies each packet without increasing the size of each packet, to obtain network-wide metrics without MTU modification, 2) allocates each packet an area in memory, called backpack, and carries metadata in it to largely reduce bandwidth overhead. vNetRadar is implemented based on the extended Berkeley Packet Filter (eBPF) and is mainly in kernel space, avoiding the CPU overhead of copying packets to user space when performing the fine-grained measurement. Evaluation results show that the easy-to-deploy vNetRadar can get fine-grained network-wide metrics with low CPU and bandwidth overhead. Tie Ma, Jin Zhang 0001, Long Luo, Hong-Fang Yu, Gang Sun 0001, Jian Sun 0019 |
GLOBECOM | 3 |
| 2022 | Fast Parameter Synchronization for Distributed Learning with Selective MulticastabstractRecent advances in distributed machine learning show theoretically and empirically that, for many models, provided workers would participate in the synchronizations eventually, i) the training still converges, even if only p workers take part in each round of synchronization, and ii) a larger p generally leads to a faster rate of convergence. These findings shed light on eliminating the bottleneck effects of parameter synchronization in large-scale data-parallel distributed training, having motivated several optimization designs.In this paper, we focus on optimizing the parameter synchronization for peer-to-peer distributed learning, in which workers generally broadcast or multicast their updated parameters to others for synchronization, and propose SELMCAST, an expressive and Pareto-optimal multicast receiver selection algorithm, to achieve the goal. Compared with the state-of-the-art design that randomly selects exactly p receivers for each worker’s multicast in a bandwidth-agnostic way, SELMCAST chooses receivers based on the global view of their available bandwidth and loads, yielding two advantages. Firstly, it could optimize the bottleneck sending rate, thus cutting down the time cost of parameter synchronization. Secondly, when more than p receivers are with sufficient bandwidth, they would be selected as many as possible, bringing benefits to the convergence of training. Extensive evaluations show that SELMCAST is efficient and always achieves near-optimal performance. Shouxi Luo, Pingzhi Fan, Ke Li 0020, Huanlai Xing, Long Luo, Hong-Fang Yu |
ICC | 5 |
| 2022 | Eliminating Communication Bottlenecks in Cross-Device Federated Learning with In-Network Processing at the EdgeabstractNowadays, cross-device federated learning (FL) is the key to achieving personalization services for mobile users and has been widely employed by companies like Google, Microsoft, and Alibaba in production. With the explosive increase of participants, the central FL server, which acts as the manager and aggregator of cross-device model training, would get overloaded, becoming the system bottlenecks. Inspired by the emerging wave of edge computing, an interesting question is: could edge clouds help cross-device FL systems overcome the bottleneck?This article provides a cautiously optimistic answer by proposing INP, an FL-specific In-Network Processing framework, along with the novel Model Download Protocol of MDP and Model Upload Protocol of MUP. With MDP and MUP, edge cloud nodes along the paths in INP can easily eliminate duplicated model downloads and pre-aggregate associated gradient uploads for the central FL server, thus alleviating its bottleneck effect, and further accelerating the entire training progress significantly. Shouxi Luo, Pingzhi Fan, Huanlai Xing, Long Luo, Hong-Fang Yu |
ICC | 4 |
| 2022 | Flexible and Efficient Multicast Transfers in Inter-Datacenter NetworksabstractThe explosive growth of global distributed services has led to a massive increase in bulk multicast data transfers over the inter-datacenter Wide-Area Network. While many solutions have been proposed to improve the performance of inter-DC bulk data transfers, they are insufficient to optimize multicast transfers because they fail to explore the characteristics of multicast transfers and network topology. This paper presents FlexCast, a flexible and efficient solution to optimize the completion times for multicast transfers. FlexCast takes advantage of topological characteristics to divide network sites into groups, partition receivers into subsets, and construct load-adaptive Steiner trees for receiver partitions to reduce completion time. It also employs a flexible multicast model for parallel transmission. For better performance FlexCast uses multiple scheduling policies to handle offline request submission, and for greater efficiency it adopts a combination of small-scale optimization and fast heuristic to address online request submission quickly. Simulations on real-world topologies show that FlexCast improves the completion time for multicast receivers by up to 80% compared to prior solutions. Long Luo, Linjian Yu, Tie Ma, Hong-Fang Yu |
IWQoS | 1 |
| 2022 | Joint Client Selection and Resource Allocation for Federated Learning in Mobile Edge NetworksabstractFederated Learning (FL) has received widespread attention in 5G mobile edge networks (MENs) due to its ability to facilitate collaborative learning of machine learning models without revealing user privacy data. However, FL training is both time and energy consuming. Constrained by the instability and limited resources of clients in MENs, it is challenging to optimize both learning time and energy consumption for FL. This paper studies the problem of client selection and resource allocation to minimize the energy consumption and learning time of multiple FL jobs competing for resources. Because minimizing learning time and minimizing energy consumption are conflicting objectives, we design a decoupling algorithm to optimize them separately and efficiently. Simulations based on popular models and learning datasets show the effectiveness of our approach, reducing up to 75.7% energy consumption and 38.5% learning time compared to prior work. Long Luo, Qingqing Cai, Zonghang Li, Hong-Fang Yu |
WCNC | 1 |
| 2022 | Blockchain-Enabled Two-Way Auction Mechanism for Electricity Trading in Internet of Electric VehiclesabstractAs people pay more attention to environment protection, the number of electric vehicles (EVs) is gradually increasing. Energy trading management for EVs is becoming a challenge. However, existing research has not considered the problem of information sharing between energy traders and issues surrounding the protection of user privacy. Therefore, in this article, we propose a vehicle-to-vehicle (V2V) and vehicle-to-grid (V2G) electricity trading architecture based on blockchain. All energy transactions of EVs can be recorded on the blockchain ledger to ensure privacy and smart contracts work as agents for pricing and optimal energy allocation. Furthermore, we introduce a two-way auction mechanism based on the Bayesian game and design a new price adjustment strategy. Finally, we propose a bidirectional auction mechanism based on the Bayesian game approach. We use extensive simulations to evaluate the performance of our proposed algorithm. Simulation results show that the social welfare and cost performance of our algorithm can be improved by up to 102.8% and 319%, respectively. Long Luo, Jingcui Feng, Hong-Fang Yu, Gang Sun 0001 |
IEEE Internet Things J. | 1 |
| 2022 | Optimizing multicast flows in high-bandwidth reconfigurable datacenter networks
Long Luo, Klaus-Tycho Förster, Stefan Schmid 0001, Hong-Fang Yu |
J. Netw. Comput. Appl. | 1 |
| 2022 | Deadline-Aware Fast One-to-Many Bulk Transfers over Inter-Datacenter NetworksabstractAn increasing number of cloud services are operated globally, where the service data are frequently replicated across geographically distributed datacenters to improve service quality and reliability. Such replication generates many one-to-many bulk data transfers over inter-datacenter networks from one datacenter to many receiver datacenters. To provide end-users with guaranteed services, these data transfers are usually required to be completed within designated deadlines. Despite the exponential growth in data demand, there has been little work on guaranteeing deadlines for one-to-many transfers, which is the subject of this paper. This paper proposes a centralized admission control coupled with a scheduling algorithm, named deAdline-Guaranteed transfEr (AGE), to guarantee the deadline of admitted data transfers and utilize the network capacity efficiently. The key idea is to flexibly select the source datacenter for receiver datacenters and allow the remaining receivers to obtain a replica from either the original source or the other receivers that have already received a copy. By jointly allocating the source for receivers and the bandwidth and routing paths for every data transfer, AGE maximizes the number of deadline-satisfied transfers. Our simulations show that compared to the state-of-the-art, AGE guarantees the deadline for up to 70 percent more transfers, achieves at least 2× higher network throughput, and reduces the completion time up to 80 percent. Long Luo, Yijing Kong, Mohammad Noormohammadpour, Zilong Ye, Gang Sun 0001, Hong-Fang Yu, Bo Li 0001 |
IEEE Trans. Cloud Comput. | 1 |
| 2022 | Intersection-Based V2X Routing via Reinforcement Learning in Vehicular Ad Hoc NetworksabstractWith the rapid development of the Internet of vehicles (IoV), routing in vehicular ad hoc networks (VANETs) has become a popular research topic. Due to the features of the dynamic network structure, constraints of road topology and variable states of vehicle nodes, VANET routing protocols face many challenges, including intermittent connectivity, large delay and high communication overhead. Location-based geographic routing is the most suitable method for VANETs, and such routing performs well on paths with an appropriate vehicle density and network load. We propose an intersection-based V2X routing protocol that includes a learning routing strategy based on historical traffic flows via Q-learning and monitoring real-time network status. The hierarchical routing protocol consists of two parts: a multidimensional Q-table, which is established to select the optimal road segments for packet forwarding at intersections; and an improved greedy strategy, which is implemented to select the optimal relays on paths. The monitoring models can detect network load and adjust routing decisions in a timely manner to prevent network congestion. This method minimizes the communication overhead and latency and ensures reliable transmission of packets. We compare our algorithm with three benchmark algorithms in an extensive simulation. The results show that our algorithm outperforms the existing methods in terms of network performance, including packet delivery ratio, end-to-end delay, and communication overhead. Long Luo, Hong-Fang Yu, Gang Sun 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Game Theoretic Approach for Multipriority Data Transmission in 5G Vehicular NetworksabstractThe vehicle-to-vehicle (V2V) communication driven by the fifth generation (5G) cellular mobile network with the features of ultra-high reliability and low latency provides promising solutions to various applications in the intelligent transportation system (ITS). To improve the resource utilization and guarantee the quality-of-service (QoS), users in 5G vehicular networks have to select appropriate communication modes and control their own transmission power. However, the highly dynamic network topology and channel status pose challenges to the mode selection. In this paper, we propose a scheme for joint mode selection and power adaptation based on the game theoretic approach with the objective of maximizing the overall system throughput. We consider the transmission requirements of multi-priority packets of different vehicular applications, where packets with higher priority have more stringent latency constraints. The segmented auction method with reserve price is performed to select modes for the vehicular users (VUEs) and the Stackelberg gaming model is introduced to solve the problem of cochannel interference. We compare our approach with three existing methods in extensive simulations. The results show that our approach outperforms the existing methods in terms of network performance, including the network throughput, resource utilization and QoS violation rate. Gang Sun 0001, Long Luo, Hong-Fang Yu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | DGT: A contribution-aware differential gradient transmission mechanism for distributed machine learning
Huaman Zhou, Zonghang Li, Qingqing Cai, Hong-Fang Yu, Shouxi Luo, Long Luo, Gang Sun 0001 |
Future Gener. Comput. Syst. | 6 |
| 2021 | Latency performance modeling and analysis for hyperledger fabric blockchain network
Xiaoqiong Xu, Gang Sun 0001, Long Luo, Huilong Cao, Hong-Fang Yu, Athanasios V. Vasilakos |
Inf. Process. Manag. | 3 |
| 2021 | Parallel spiral search algorithm applied to integer motion estimation
Longzhao Shi, Long Luo, Xiuzhi Yang |
Signal Process. Image Commun. | 3 |
| 2021 | TSEngine: Enable Efficient Communication Overlay in Distributed Machine Learning in WANsabstractIn recent years, distributed machine learning in WANs (DML-WANs), i.e., collaboratively training a high-quality ML model cross geo-distributed micro-clouds or edge devices, has attracted attention and been widely applied. Compared with cloud-centric training, DML-WANs avoids the high cost of transferring large amounts of raw data to a central cloud and privacy concerns. However, performing DML-WANs still faces challenges. Model synchronization, an essential step of DML-WANs, is accompanied by a lot of model communication cross limited-bandwidth WANs, which generates high communication overhead. Moreover, the parameter server system, which has been widely used, performs model synchronization in a centralized manner, resulting in serious communication in-cast problem. Such communication in-cast further raises the communication overhead, leading to the low efficiency of DML-WANs. To alleviate the communication in-cast, existing researches attempt to build tree-based communication overlays over the parameter server and workers. However, we identify that these approaches can not adapt to the dynamic and heterogeneous network of DML-WANs, resulting in insufficient improvements. This paper proposes TSEngine, an adaptive communication scheduler for efficient communication overlay of the parameter server system in DML-WANs. Its core idea is to dynamically schedule the communication logic over the parameter server and workers based on the active network perception. Specifically, we propose novel communication scheduling protocols for model distribution and model aggregation, respectively. We have implemented TSEngine in a mainstream parameter server system and verified its effectiveness in DML-WANs testbeds. Huaman Zhou, Weibo Cai, Zonghang Li, Hong-Fang Yu, Long Luo, Gang Sun 0001 |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2020 | Migration Aware Virtual Network Function Placing and Routing in Uncertain EnvironmentabstractNetwork Function Virtualization (NFV) aims to provide a way to build agile and service-aware networks by building a new paradigm of provisioning network services where physical network functions are deployed as Virtual Network Functions (VNFs). However, how to optimally allocate resources in the uncertain NFV environment where flow rates fluctuate has not been fully resolved. In this paper, we study the cost minimizing problem of VNF placing and routing optimization while two goals are considered: optimizing service migration caused by flow fluctuation and stabilizing queue backlogs in the network. We first formulate the problem as a stochastic optimization programming problem. Then we propose an online algorithm, named Migration Aware VNF plaCing and Routing Online algorithm (MACRO), based on Lyapunov optimization technique. MACRO can make good decisions without knowing any future information. The theoretical analysis suggests that MACRO achieves an optimality gap of O([1/(V)]) and the queue backlogs are bounded by O(V), where V is a tunable parameter that controls the tradeoff between cost and backlogs. The experiment results suggest that MACRO achieves queue stability and outperforms benchmark algorithm by 6% in terms of cost. Yanghao Xie, Sheng Wang 0006, Long Luo |
GLOBECOM | 4 |
| 2020 | Virtualized Network Function Provisioning in Stochastic Cloud EnvironmentabstractNetwork Function Virtualization (NFV) provides a new paradigm for provisioning network service where network functions are deployed as Virtual Network Functions (VNFs). Due to the advantages of NFV, many Network Function Virtualization Providers (NFVPs) offer their NFV services by deploying VNFs with purchased cloud resources in cloud environment to save the provisioning expense. However, existing VNF provisioning solutions ignore the influences of the dynamics of cloud environment, which may lead to over-provisioning and high deployment expense. In this paper, we study the problem of how should the NFVPs purchase cloud resources to provide NFV services for customers in order to minimize the expense of NFVPs, considering the dynamics of the system. We first abstract the system model of this problem and formulate it as a stochastic optimization programming problem. Then, we present our VIrtual Network functiOn proviSioning (VINOS) approach that can efficiently solve the stochastic optimization programming with a rolling horizon procedure. In particular, it first leverages Long Short Term Memory (LSTM) networks to predict future exogenous information and then optimally solves a deterministic problem over short horizon. We conduct extensive numerical experiments to evaluate the proposed approach. The experiment results suggest that our approach achieves total cost of 1.2 times offline optimum, and outperforms the benchmark algorithm by 8%, averagely. Yanghao Xie, Sheng Wang 0006, Long Luo |
ICC | 4 |
| 2020 | SplitCast: Optimizing Multicast Flows in Reconfigurable Datacenter NetworksabstractMany modern cloud applications frequently generate multicast traffic, which is becoming one of the primary communication patterns in datacenters. Emerging reconfigurable datacenter technologies enable interesting new opportunities to support such multicast traffic in the physical layer: novel circuit switches offer high-performance inter-rack multicast capabilities. However, not much is known today about the algorithmic challenges introduced by this new technology.This paper presents SplitCast, a preemptive multicast scheduling approach that fully exploits emerging physical-layer multicast capabilities to reduce flow times. SplitCast dynamically reconfigures the circuit switches to adapt to the multicast traffic, accounting for reconfiguration delays. In particular, SplitCast relies on simple single-hop routing and leverages flexibilities by supporting splittable multicast so that a transfer can already be delivered to just a subset of receivers when the circuit capacity is insufficient. Our evaluation results show that SplitCast can reduce flow times significantly compared to state-of-the-art solutions. Long Luo, Klaus-Tycho Förster, Stefan Schmid 0001, Hong-Fang Yu |
INFOCOM | 1 |
| 2020 | Job scheduling for distributed machine learning in optical WAN
Hong-Fang Yu, Gang Sun 0001, Long Luo, Qixuan Jin, Shouxi Luo |
Future Gener. Comput. Syst. | 4 |
| 2020 | Deadline-Aware Multicast Transfers in Software-Defined Optical Wide-Area NetworksabstractThe increasing amount of data replication across datacenters introduces a need for efficient bulk data transfer protocols which provide certain guarantees, most notably timely transfer completion. We present DaRTree which leverages emerging optical reconfiguration technologies, to jointly optimize topology and multicast transfers in software-defined optical Wide-Area Networks (WANs), and thereby maximize throughput and acceptance ratio of transfer requests subject to transfer deadlines. DaRTree is based on a novel integer linear program relaxation and deterministic rounding scheme. To this end, DaRTree uses Steiner trees for forwarding and adaptive routing based on the current network load. DaRTree provides transfer completion guarantees without the need for rescheduling or preemption. Our evaluations show that DaRTree increases the network throughput and the number of accepted requests by up to 1.7×, especially for larger WANs. Moreover, DaRTree even outperforms state-of-the-art solutions when the traffic demands are only unicast transfers or when the WAN topology cannot be reconfigured. While DaRTree determines the rate and route to serve a request at the time of (online) admission control, we show that the acceptance ratio and throughput can be improved by up to 1.3× even further when DaRTree updates the rate and route of admitted transfers also at runtime. Long Luo, Klaus-Tycho Förster, Stefan Schmid 0001, Hong-Fang Yu |
IEEE J. Sel. Areas Commun. | 1 |
| 2019 | DaRTree: deadline-aware multicast transfers in reconfigurable wide-area networksabstractThe increasing amount of data replication across datacenters introduces a need for efficient bulk data transfer protocols which meet QoS guarantees, notably timely completion. We present DaRTree which leverages emerging optical reconfiguration technologies, to jointly optimize topology and multicast transfers, and thereby maximize throughput and acceptance ratio of transfer requests subject to deadlines. DaRTree is based on a novel integer linear program relaxation and deterministic rounding scheme. To this end, DaRTree uses multicast Steiner trees and adaptive routing based on the current network load. DaRTree provides its guarantees without need for rescheduling or preemption. Our evaluations show that DaRTree increases the network throughput and the number of accepted requests by up to 70%, especially for larger Wide-Area Networks (WANs). In fact, we also find that DaRTree even outperforms state-of-the-art solutions when the network scheduler is only capable of routing unicast transfers or when the WAN topology is bound to be non-reconfigurable. Long Luo, Klaus-Tycho Förster, Stefan Schmid 0001, Hong-Fang Yu |
IWQoS | 1 |
| 2019 | Customizable network update planning in SDN
Shouxi Luo, Hong-Fang Yu, Long Luo, Lemin Li |
J. Netw. Comput. Appl. | 3 |
| 2019 | Scalable explicit path control in software-defined networks
Long Luo, Hong-Fang Yu, Shouxi Luo, Zilong Ye, Xiaojiang Du, Mohsen Guizani |
J. Netw. Comput. Appl. | 1 |
| 2018 | A Multi-model Fusion Framework based on Deep Learning for Sentiment ClassificationabstractWith the development of the Internet, more and more data can be found on texts information. People produce texts information via writing blogs, product reviews, microblogs, film reviews and so on, which contain sentiments or opinions of the writer. User comments usually can reflect their intuitive feelings for a product. We can dig out relatively large value through sentiment analysis for these comments. Sentiment classification, as one of the most important tasks in sentiment analysis for many real world applications, is our main focus in this article. To improve the accuracy of sentiment classification, we propose a deep neural network fusion framework, which is composed of a multi-window CNN-LSTM model and a multi-window CNN-CNN model with the fusion of the probability of two models to generate final output. Experimental results instruct that our framework is quite feasible. Fen Yang, Jia Zhu 0003, Xuming Wang, Xingcheng Wu, Yong Tang 0001, Long Luo |
CSCWD | 6 |
| 2018 | Deadline-Guaranteed Point-to-Multipoint Bulk Transfers in Inter-Datacenter NetworksabstractMany modern cloud services are operated across geographically distributed datacenters, and they are usually associated with a demand of transferring bulk data among datacenters for achieving a high performance and reliability. These transfers (e.g., data replications and synchronizations) may require the inter-datacenter networks to deliver data from one Point (or datacenter) to Multiple Points (or datacenters) and impose a deadline on the transfer for providing a guaranteed service to end users. However, very little work has been done to manage an efficient point-to- multipoint bulk transfer while considering the deadline requirements. In this paper, we investigate the deadline- aware point-to-multipoint (P2MP) transfer problem and propose a centralized deAdline- Guaranteed transfEr (AGE) approach that can guarantee the deadline for P2MP transfers while efficiently utilizing the inter-datacenter bandwidth resources. For each arriving request, AGE jointly determines the transfer source selections and bandwidth allocations such that the number of deadline-guaranteed transfers can be maximized. Our simulation experiments show that AGE can accommodate up to 53.3% more transfers whose deadline requirements are met. In addition, AGE can achieve 20% higher network utilization than prior bulk transfer approaches. Long Luo, Hong-Fang Yu, Zilong Ye |
ICC | 1 |
| 2018 | Online Deadline-Aware Bulk Transfer Over Inter-Datacenter WANsabstractMany large-scale compute-intensive and mission-critical online service applications are being deployed on geo-distributed datacenters, which require transfers of bulk business data over Wide Area Networks (WANs). The bulk transfers are often associated with different requirements on deadlines, either a complete transfer before a hard deadline or a best-effort delivery within a soft deadline. In this paper, we study the online bulk transfer problem over inter-datacenter WANs, while taking into consideration the requests with a mixture of hard and soft deadlines. We use Linear Programming (LP) to mathematically formulate the problem with the objective of maximizing a system utility represented by the service provider's revenue, taking into account the revenue earned from deadline-met transfers and the penalty paid for deadline-missed ones. We propose an online framework to efficiently manage mixed bulk transfers and design a competitive algorithm that applies the primal-dual method to make routing and resource allocation based on the LP. We perform theoretical analysis to prove that the proposed approach can achieve a competitive ratio of (e-1)/e with little link capacity augmentation. In addition, we conduct comprehensive simulations to evaluate the performance of our method. Simulation results show that our method irrespective of the revenue model, can accept at least 25% more transfer requests and improve the network utilization by at least 35%, compared to prior solutions. Long Luo, Hong-Fang Yu, Zilong Ye, Xiaojiang Du |
INFOCOM | 1 |
| 2016 | Achieving Fast and Lightweight SDN Updates with Segment RoutingabstractIn SDN, forwarding rules are frequently updated to adapt to network dynamics. During the procedure, path consistency needs to be preserved; otherwise, in-flight packets might meet with forwarding errors such as loops and black holes. Despite a large number of suggestions have been proposed, they take either a long duration or have high rule-space overheads, thus fail to be practical for large-scale high dynamic networks. In this paper, we propose FLUS, a Segment Routing (SR) based mechanism, to achieve fast and lightweight path updates. Basically, when a route needs a change, FLUS instantly employs SR to construct its desired new path by concatenating some fragments of the already existing paths. After the actual paths are established, FLUS then shifts incoming packets to them and disables the transitional ones. Such a design helps packets enjoy their new paths immediately without introducing rule-space overheads. This paper presents FLUS's segment allocation, path construction, and the corresponding optimal algorithms in detail. Our evaluation based on real and synthesized networks shows: FLUS can handle up to 92-100% updates using SR in real-time and save 72-88% rule overhead compared to prior methods. Long Luo, Hong-Fang Yu, Shouxi Luo, Mingui Zhang, Shui Yu 0001 |
GLOBECOM | 1 |
| 2015 | Fast lossless traffic migration for SDN updatesabstractMigration of traffic from one configuration to another is common in SDNs due to node/link failures, network maintenance, policy reconfiguration, intrusion detection, network upgrades, and etc. When the network devices are informed by the controller to execute the traffic migration, it's difficult even impossible to force all network devices to perform the update action in a strict synchronized way. Thus the network is likely to see transient overlapped traffic from both the new configuration and the old one. This kind of overlap may cause overload to those hot spots. This paper reveals the transient congestion problem during traffic migration in an SDN update. According to the observation, it's feasible for the controller to schedule ingress nodes to perform the migration in an order thus the transient congestion is avoided. This scheduling problem is formulated as a Mixed Integer Linear Program (MIP) model. If feasible orders exist for ingress network nodes to perform the migration, the MIP can always find the order that achieves the minimum steps in all possibilities. A heuristic method (named ATOMIP (ATOmic-MIP)) is proposed to speed-up the solving of this MIP. Evaluation based on network topologies observed from real ISPs shows that the lossless migration happens in sub-seconds. Long Luo, Hong-Fang Yu, Shouxi Luo, Mingui Zhang |
ICC | 1 |