VLDB 2026 Research / reviewers in the wild / expert
Zonghang Li
dblp:218/5224
· DBLP profile ↗
21ranked-venue papers
4as first author
20since 2021 · last 2025
0000-0002-2796-039XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 2 first-author · 11 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Which Cluster Meets My Deadline: a Budget-Aware Scheduler for Distributed Training Jobs in Heterogeneous EnvironmentsabstractTraining deep learning (DL) models demands substantial computational resources, often relying on expensive GPUs in a distributed manner. To meet this demand, cloud providers deploy GPU clusters worldwide to offer users compute instance rental services. These GPU clusters are typically heterogeneous, comprising multiple GPU types with varying computational capabilities, and their prices vary significantly across both GPU instance types and geographic regions. Meanwhile, users often have specific deadlines for training their models. Most existing works focus solely on performance, overlooking price heterogeneity and failing to optimize costs effectively. Given the high cost and time demands of model training, focusing on performance alone is insufficient. In this paper, we aim to balance performance and cost through a DL broker service that maximizes the number of jobs completed within their deadlines under a given budget. We propose CADDS, which jointly optimizes cluster placement and dynamically adjusts GPU type and quantity during training to reduce rental costs. We formulate the scheduling problem as an integer nonlinear programming problem and propose an efficient online approach combining greedy and dynamic programming. Experimental results demonstrate that CADDS significantly outperforms existing approaches, improving the deadline satisfactory ratio within budget limits by$\mathbf{5 9. 4 \%}$to$\mathbf{8 0. 2 \%}$. Long Luo, Zonghang Li, Gang Sun 0001, Hong-Fang Yu |
ICC | 3 |
| 2025 | TopoDT: Digital Twin-Assisted UAV Topology Optimization for Targets TrackingabstractUnmanned Aerial Vehicles (UAVs) have been an attractive device to serve target tracking scenarios, such as hit- and-run tracking and border patrol. Nonetheless, it is difficult to implement real-time UAV topology optimization due to communication resources of UAVs and random moving speeds of targets. To address the problem, we propose a Digital Twins-assisted topology optimization framework (TopoDT). We formulate a UAV topology optimization model based on Lyapunov theory in the framework. The model is decoupled into two subproblems using our proposed TopoDT topology optimization algorithm. The DT model can allow UAVs to implement neighbor selection to construct and optimize small-scale local topologies for tracking low-speed moving targets. In addition, it allows UAVs to construct large-scale global topologies for tracking high-speed moving targets based on trajectory derivation. The system simulation results demonstrate that our solution reduces the end-to-end latency by 63.0% while decreasing the hop counts by 50% compared to state-of-the-art benchmarks. Longyu Zhou, Supeng Leng, Zonghang Li, Tony Q. S. Quek |
IWCMC | 3 |
| 2025 | TPI-LLM: Serving 70B-Scale LLMs Efficiently on Low-Resource Mobile DevicesabstractLLM serving is shifting from cloud to edge due to privacy concerns over user interaction data. However, mobile devices struggle with very limited computing power and memory, requiring collaboration among multiple devices to run LLM apps. The mainstream solution, pipeline parallelism, is inefficient for such cases because mobile devices typically run only one inference task at a time. This article argues that tensor parallelism, despite its high communication cost, can better fit such scenarios. We introduce TPI-LLM, a compute and memory-efficient tensor parallel inference system designed to run 70B-scale LLMs on low-resource mobile devices. It keeps sensitive raw data local on users’ devices and employs a sliding window memory scheduler to dynamically manage layer weights. It overlaps disk I/O with computation and communication, enabling efficient operation of large models on memory-limited devices. Extensive experiments show that TPI-LLM reduces token latency by 80%–90% compared to Transformers, Accelerate, and Galaxy. It also cuts the peak memory footprint by 90%, requiring just 3.1 GiB of memory for 70B-scale models. Zonghang Li, Wenjiao Feng, Mohsen Guizani, Hong-Fang Yu |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | Collaborative Multimodal Vehicular Transformer Training Using Federated LearningabstractThe Internet of Vehicles (IoV) is an intricate ecosystem brimming with diverse data modalities, including visual streams from cameras, GPS-based location information, sensor-derived operational metrics, and auditory commands from users. These necessitate advanced multimodal learning capabilities. The Transformer architecture, a significant innovation in artificial intelligence, has demonstrated its proficiency in representing varied modalities, facilitating multimodal machine learning. However, its direct application within the IoV is hampered by concerns over user data privacy. Federated learning (FL), a distributed learning paradigm, offers a solution that upholds data privacy. We propose a novel multimodal Transformer-based federated learning framework that capitalizes on the Transformer's ability to effectively handle multimodal data, enabling collaborative learning across heterogeneous data sources. This framework aligns with stringent privacy regulations, enhancing the protection of user data privacy while also boosting learning efficiency. Our approach surpasses traditional models in both accuracy and efficiency, presenting a significant advance in the multimodal machine learning domain. Xingjian Cao, Zonghang Li, Gang Sun 0001, Hong-Fang Yu |
VTC Spring | 2 |
| 2024 | Effective Intrusion Detection in Highly Imbalanced IoT Networks With Lightweight S2CGAN-IDSabstractSince the advent of the Internet of Things (IoT), exchanging vast amounts of information has increased the number of security threats in networks. As a result, intrusion detection based on deep learning (DL) has been developed to achieve high throughput and high precision. Unlike general deep learning-based scenarios, IoT networks contain benign traffic far more than abnormal traffic, with some rare attacks. However, most existing studies have been focused on sacrificing the detection rate of the majority class in order to improve the detection rate of the minority class in class-imbalanced IoT networks. Although this way can reduce the false negative rate of minority classes, it both wastes resources and reduces the credibility of the intrusion detection systems. To address this issue, we propose a lightweight framework named S2CGAN-IDS. The proposed framework leverages the distribution characteristics of network traffic to expand the number of minority categories in both data space and feature space, resulting in a substantial increase in the detection rate of minority categories while simultaneously ensuring the detection precision of majority categories. To reduce the impact of sparsity on the experiments, the CICIDS2017 numeric dataset is utilized to demonstrate the effectiveness of the proposed method. The experimental results indicate that our proposed approach outperforms the superior method in both Precision and Recall, particularly with a 10.2% improvement in the F1-score. Caihong Wang, Du Xu, Zonghang Li, Dusit Niyato |
IEEE Internet Things J. | 3 |
| 2024 | Energy-Efficient Hierarchical Collaborative Learning Over LEO Satellite ConstellationsabstractThe hierarchical collaborative learning within Low Earth Orbit (LEO) satellite constellations, termed LEO-HCL, is gaining increasing popularity by integrating intra-orbit Inter-Satellite Links and orbital edge computing to alleviate the latency issues caused by intermittent satellite connectivity in satellite-ground training architectures. However, LEO-HCL systems are confronted with a triad of challenges: the variable topology induced by satellite mobility, limited onboard computing and communication resources, and stringent energy constraints. In response to these challenges, we propose an energy-efficient training algorithm called FedAAC, which adaptively optimizes both aggregation frequency and model compression ratio within the resource-constrained LEO network. We have conducted a theoretical analysis of model convergence and investigated the relationship between convergence, aggregation frequency, and model compression ratio. Building on this analysis, we offer an approximation algorithm that dynamically calculates the optimal aggregation frequency and compression ratio during the training process. Extensive simulations have demonstrated that FedAAC significantly outperforms existing methods, offering enhanced convergence speed and energy efficiency. Compared to prior solutions, FedAAC achieves a 60% reduction in energy consumption, a 70% decrease in training time, and a 52% lower communication overhead. Long Luo, Chi Zhang 0076, Hong-Fang Yu, Zonghang Li, Gang Sun 0001, Shouxi Luo |
IEEE J. Sel. Areas Commun. | 4 |
| 2024 | Diffusion-Based Reinforcement Learning for Edge-Enabled AI-Generated Content ServicesabstractAs Metaverse emerges as the next-generation Internet paradigm, the ability to efficiently generate content is paramount. AI-Generated Content (AIGC) emerges as a key solution, yet the resource-intensive nature of large Generative AI (GAI) models presents challenges. To address this issue, we introduce an AIGC-as-a-Service (AaaS) architecture, which deploys AIGC models in wireless edge networks to ensure broad AIGC services accessibility for Metaverse users. Nonetheless, an important aspect of providing personalized user experiences requires carefully selecting AIGC Service Providers (ASPs) capable of effectively executing user tasks, which is complicated by environmental uncertainty and variability. Addressing this gap in current research, we introduce the AI-Generated Optimal Decision (AGOD) algorithm, a diffusion model-based approach for generating the optimal ASP selection decisions. Integrating AGOD with Deep Reinforcement Learning (DRL), we develop the Deep Diffusion Soft Actor-Critic (D2SAC) algorithm, enhancing the efficiency and effectiveness of ASP selection. Our comprehensive experiments demonstrate that D2SAC outperforms seven leading DRL algorithms. Furthermore, the proposed AGOD algorithm has the potential for extension to various optimization problems in wireless networks, positioning it as a promising approach for future research on AIGC-driven services. The implementation of our proposed method is available at:https://github.com/Lizonghang/AGOD. Hongyang Du 0001, Zonghang Li, Dusit Niyato, Jiawen Kang 0001, Zehui Xiong, Huawei Huang, Shiwen Mao |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | Accelerating Geo-Distributed Machine Learning With Network-Aware Adaptive Tree and Auxiliary RouteabstractDistributed machine learning is becoming increasingly popular for geo-distributed data analytics, facilitating the collaborative analysis of data scattered across data centers in different regions. This paradigm eliminates the need for centralizing sensitive raw data in one location but faces the significant challenge of high parameter synchronization delays, which stems from the constraints of bandwidth-limited, heterogeneous, and fluctuating wide-area networks. Prior research has focused on optimizing the synchronization topology, evolving from starlike to tree-based structures. However, these solutions typically depend on regular tree structures and lack an adequate topology metric, resulting in limited improvements. This paper proposes NetStorm, an adaptive and highly efficient communication scheduler designed to speed up parameter synchronization across geo-distributed data centers. First, it establishes an effective metric for optimizing a multi-root FAPT synchronization topology. Second, a network awareness module is developed to acquire network knowledge, aiding in topology decisions. Third, a multipath auxiliary transmission mechanism is introduced to enhance network awareness and facilitate multipath transmissions. Lastly, we design policy consistency protocols to guarantee seamless updates of transmission policies. Empirical results demonstrate that NetStorm significantly outperforms distributed training systems like MXNET, MLNET, and TSEngine, with a speedup of 6.5~9.2 times over MXNET. Zonghang Li, Wenjiao Feng, Weibo Cai, Hong-Fang Yu, Long Luo, Gang Sun 0001, Hongyang Du 0001, Dusit Niyato |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | Graph Learning Enhanced UAV Swarms Based Multiple Targets TrackingabstractWith the development of Artificial Intelligence (AI) technology, diverse Internet of Things (IoT) devices digesting abundant data have been exploited to meet more application requirements. In this regard, Unmanned Aerial Vehicle-based Multiple Targets Tracking (UAV-MTT) applications have been paid attention to processing a vast amount of sensing information for accurate and consecutive MTT. However, this application exposes imperative computing requirements on resource-limited UAVs. Edge computing can provide extra resources to alleviate the computing pressure for high-efficiency tracking decisions. Nonetheless, it is challenging to dynamically allocate UAVs for optimal association with time-varying target trajectories. To address the mentioned problems, we propose a terminal-edge cooperative tracking framework with a cross-layer resource cooperation method. In this design, we propose an auction-based cooperative game algorithm to implement highly accurate trajectory prediction. We then propose a graph learning-based tracking algorithm to adaptively manage the dynamic UAV topology for consecutive MTT. Simulation results demonstrate that our algorithm improves 70% prediction accuracy compared to other benchmarks while saving 40% energy consumption. Longyu Zhou, Supeng Leng, Zonghang Li, Hongyang Du 0001, Dusit Niyato |
GLOBECOM | 3 |
| 2023 | Personalized Saliency in Task-Oriented Semantic Communications: Image Transmission and Performance AnalysisabstractSemantic communication, as a promising technology, has emerged to break through the Shannon limit, which is envisioned as the key enabler and fundamental paradigm for future 6G networks and applications, e.g., smart healthcare. In this paper, we focus on UAV image-sensing-driven task-oriented semantic communications scenarios. The majority of existing work has focused on designing advanced algorithms for high-performance semantic communication. However, the challenges, such as energy-hungry and efficiency-limited image retrieval manner, and semantic encoding without considering user personality, have not been explored yet. These challenges have hindered the widespread adoption of semantic communication. To address the above challenges, at the semantic level, we first design an energy-efficient task-oriented semantic communication framework with a triple-based scene graph for image information. We then design a new personalized semantic encoder based on user interests to meet the requirements of personalized saliency. Moreover, at the communication level, we study the effects of dynamic wireless fading channel on semantic transmission mathematically and thus design an optimal multi-user resource allocation scheme by using game theory. Numerical results based on real-world datasets clearly indicate that the proposed framework and schemes significantly enhance the personalization and anti-interference performance of semantic communication, and are also efficient to improve the communication quality of semantic communication services. Jiawen Kang 0001, Hongyang Du 0001, Zonghang Li, Zehui Xiong, Shiyao Ma, Dusit Niyato |
IEEE J. Sel. Areas Commun. | 3 |
| 2023 | Cross-silo heterogeneous model federated multitask learning
Xingjian Cao, Zonghang Li, Gang Sun 0001, Hong-Fang Yu, Mohsen Guizani |
Knowl. Based Syst. | 2 |
| 2023 | HFedMS: Heterogeneous Federated Learning With Memorable Data Semantics in Industrial MetaverseabstractFederated Learning (FL), as a rapidly evolving privacy-preserving collaborative machine learning paradigm, is a promising approach to enable edge intelligence in the emerging Industrial Metaverse. Even though many successful use cases have proved the feasibility of FL in theory, in the industrial practice of Metaverse, the problems of non-independent and identically distributed (non-i.i.d.) data, learning forgetting caused by streaming industrial data, and scarce communication bandwidth remain key barriers to realize practical FL. Facing the above three challenges simultaneously, this paper presents a high-performance and efficient system namedHFedMSfor incorporating practical FL into Industrial Metaverse.HFedMSreduces data heterogeneity through dynamic grouping and training mode conversion (Dynamic Sequential-to-Parallel Training, STP). Then, it compensates for the forgotten knowledge by fusing compressed historical data semantics and calibrates classifier parameters (Semantic Compression and Compensation, SCC). Finally, the network parameters of the feature extractor and classifier are synchronized in different frequencies (Layer-wise Alternative Synchronization Protocol, LASP) to reduce communication costs. These techniques make FL more adaptable to the heterogeneous streaming data continuously generated by industrial equipment, and are also more efficient in communication than traditional methods (e.g., Federated Averaging). Extensive experiments have been conducted on the streamed non-i.i.d. FEMNIST dataset using 368 simulated devices. Numerical results show thatHFedMSimproves the classification accuracy by at least 6.4% compared with 8 benchmarks and saves both the overall runtime and transfer bytes by up to 98%, proving its superiority in precision and efficiency. Shenglai Zeng, Zonghang Li, Hong-Fang Yu, Long Luo, Bo Li 0001, Dusit Niyato |
IEEE Trans. Cloud Comput. | 2 |
| 2023 | NBSync: Parallelism of Local Computing and Global Synchronization for Fast Distributed Machine Learning in WANsabstractRecently, due to privacy concerns, distributed machine learning in Wide-Area Networks (DML-WANs) attracts increasing attention and has been widely deployed to promote the widespread application of intelligence services that rely on geographically distributed data. DML-WANs is essentially performing collaboratively federated learning over a combination of servers at both edge and cloud on a large spatial scale. However, efficient model training is challenging for DML-WANs because it is blocked by the high overhead of model parameter synchronization between computing servers over WANs. The reason is that there has a sequential dependency between local model computing and global model synchronization of traditional DML-WANs training methods intrinsically producing a sequential blockage between them, e.g., FedAvg. When the computing heterogeneity and the low WAN bandwidth coexist, a long block of global model synchronization prolongs the training time and leads to low utilization of local computing. Despite many efforts on alleviating synchronization overhead with novel communication technologies and synchronization methods, they still use traditional training patterns with sequential dependency and thereby have very limited improvements, such as FedAsync and ESync. In this article, we propose NBSync, a novel training algorithm for DML-WANs, which greatly speeds up the model training by the parallelism of local computing and global synchronization. NBSync employs a well-designed pipelining scheme, which can properly relax the sequential dependency of local computing and global synchronization and process them in parallel so as to overlap their operating overhead in the time dimension. NBSync also realizes flexible, differentiated and dynamical local computing for workers to maximize the overlap ratio in dynamically heterogeneous training environments. Convergence analysis shows that the convergence rate of NBSync training process is asymptotically equal to that of SSGD, and NBSync has a better convergence efficiency. We implemented the prototype of NBSync based on a popular parameter server system, i.e., MXNET's PS-LITE library, and evaluate its performance on a DML-WANs testbed. Experimental results show that NBSync speeds up training about 1.43×–2.79× than state-of-the-art distributed training algorithms (DTAs) in DML-WANs scenarios where computing heterogeneity and low WAN bandwidth coexist. Huaman Zhou, Zonghang Li, Hong-Fang Yu, Long Luo, Gang Sun 0001 |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | Heterogeneous Federated Learning via Grouped Sequential-to-Parallel Training
Shenglai Zeng, Zonghang Li, Hong-Fang Yu, Yihong He, Zenglin Xu, Dusit Niyato, Han Yu 0001 |
DASFAA (2) | 2 |
| 2022 | Joint Client Selection and Resource Allocation for Federated Learning in Mobile Edge NetworksabstractFederated Learning (FL) has received widespread attention in 5G mobile edge networks (MENs) due to its ability to facilitate collaborative learning of machine learning models without revealing user privacy data. However, FL training is both time and energy consuming. Constrained by the instability and limited resources of clients in MENs, it is challenging to optimize both learning time and energy consumption for FL. This paper studies the problem of client selection and resource allocation to minimize the energy consumption and learning time of multiple FL jobs competing for resources. Because minimizing learning time and minimizing energy consumption are conflicting objectives, we design a decoupling algorithm to optimize them separately and efficiently. Simulations based on popular models and learning datasets show the effectiveness of our approach, reducing up to 75.7% energy consumption and 38.5% learning time compared to prior work. Long Luo, Qingqing Cai, Zonghang Li, Hong-Fang Yu |
WCNC | 3 |
| 2022 | Data Heterogeneity-Robust Federated Learning via Group Client Selection in Industrial IoTabstractNowadays, the Industrial Internet of Things (IIoT) has played an integral role in Industry 4.0 and produced massive amounts of data for industrial intelligence. These data locate on decentralized devices in modern factories. To protect the confidentiality of industrial data, federated learning (FL) was introduced to collaboratively train shared machine learning (ML) models. However, the local data collected by different devices skew in class distribution and degrade industrial FL performance. This challenge has been widely studied at the mobile edge, but they ignored the rapidly changing streaming data and clustering nature of factory devices, and more seriously, they may threaten data security. In this article, we propose FED GS, which is a hierarchical cloud-edge-end FL framework for 5G empowered industries, to improve industrial FL performance on non-independent and identically distributed (non-j) data. Taking advantage of naturally clustered factory devices, FED GS uses a gradient-based binary permutation algorithm (GBP-CS) to select a subset of devices within each factory and build homogeneous super nodes participating in FL training. Then, we propose a compound-step synchronization protocol to coordinate the training process within and among these super nodes, which shows great robustness against data heterogeneity. The proposed methods are time-efficient and can adapt to dynamic environments, without exposing confidential industrial data in risky manipulation. We prove that FED GS has better convergence performance than FedAvg and give a relaxed condition under which FED GS is more communication efficient. The extensive experiments show that FED GS improves accuracy by 3.5% and reduces training rounds by 59% on average, confirming its superior effectiveness and efficiency on non-i.i.d. data. Zonghang Li, Yihong He, Hong-Fang Yu, Jiawen Kang 0001, Xiaoping Li 0002, Zenglin Xu, Dusit Niyato |
IEEE Internet Things J. | 1 |
| 2022 | ESync: Accelerating Intra-Domain Federated Learning in Heterogeneous Data CentersabstractFederated Learning (FL) serves privacy-preserving collaborative learning among multiple isolated parties, while retaining their privacy data locally. Cross-device and cross-silo FL have achieved great success in cross-domain applications, in which the scarce communication resource is the primary bottleneck. Driven by the need to combine heterogeneous machines from different parties to build a shared data center, we foundintra-domain FL, a new type of FL in which isolated parties collaborate in the shared data center, and strong computational heterogeneity becomes the primary bottleneck. To mitigate the training inefficiency caused by stragglers, this article proposes an efficient synchronization algorithmESync, which allows parties to train different iterations locally under the coordination of a novel schedulerState Server. We give the boundaries of weight divergence and optimality gap ofESync, and analyze the trade-off between convergence accuracy and communication efficiency. Extensive experiments are conducted to compareESyncwith SSGD, ASGD, DC-ASGD, FedAvg, FedAsync, TiFL, and FedDrop under strong computational heterogeneity. Numerical results show thatESyncachieves great speed up without loss of accuracy, and therefore demonstrate the effectiveness ofESyncin both training efficiency and converged accuracy. Zonghang Li, Huaman Zhou, Tianyao Zhou, Hong-Fang Yu, Zenglin Xu, Gang Sun 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2021 | DGT: A contribution-aware differential gradient transmission mechanism for distributed machine learning
Huaman Zhou, Zonghang Li, Qingqing Cai, Hong-Fang Yu, Shouxi Luo, Long Luo, Gang Sun 0001 |
Future Gener. Comput. Syst. | 2 |
| 2021 | Mitigating Conflicting Transactions in Hyperledger Fabric-Permissioned Blockchain for Delay-Sensitive IoT ApplicationsabstractBlockchain is a promising emerging technology that is envisioned to play a key role in establishing secure and reliable Internet-of-Things (IoT) ecosystems without the involvement of any third party. Hyperledger Fabric, a permissioned blockchain system that can yield high throughput and low consensus delay, has shown its capability in enhancing security and privacy protection for delay-sensitive IoT services. The literature, however, has not considered the conflicting transaction problem which may substantially limit the system performance and degrade QoS for the end users. In this article, we propose CATP-Fabric, a new blockchain system to address the conflicting transaction problem by reducing the number of potentially conflicting transactions with less overhead. First, the transactions within a block are divided into different groups to facilitate parallel transaction processing. Then, CATP-Fabric filters stale transactions and prioritizes the read-only transactions in each group to eliminate unnecessary overhead. Finally, we formulate the selection of aborting transactions in CATP-Fabric as a binary integer-programming problem and develop a low-complexity optimization algorithm to minimize the number of aborted transactions. Illustrative results show that our proposed CATP-Fabric blockchain system achieves high throughput of successful transactions while maintaining a lower aborting transaction rate compared to the benchmark blockchain systems. Xiaoqiong Xu, Zonghang Li, Hong-Fang Yu, Gang Sun 0001, Sabita Maharjan, Yan Zhang 0002 |
IEEE Internet Things J. | 3 |
| 2021 | TSEngine: Enable Efficient Communication Overlay in Distributed Machine Learning in WANsabstractIn recent years, distributed machine learning in WANs (DML-WANs), i.e., collaboratively training a high-quality ML model cross geo-distributed micro-clouds or edge devices, has attracted attention and been widely applied. Compared with cloud-centric training, DML-WANs avoids the high cost of transferring large amounts of raw data to a central cloud and privacy concerns. However, performing DML-WANs still faces challenges. Model synchronization, an essential step of DML-WANs, is accompanied by a lot of model communication cross limited-bandwidth WANs, which generates high communication overhead. Moreover, the parameter server system, which has been widely used, performs model synchronization in a centralized manner, resulting in serious communication in-cast problem. Such communication in-cast further raises the communication overhead, leading to the low efficiency of DML-WANs. To alleviate the communication in-cast, existing researches attempt to build tree-based communication overlays over the parameter server and workers. However, we identify that these approaches can not adapt to the dynamic and heterogeneous network of DML-WANs, resulting in insufficient improvements. This paper proposes TSEngine, an adaptive communication scheduler for efficient communication overlay of the parameter server system in DML-WANs. Its core idea is to dynamically schedule the communication logic over the parameter server and workers based on the active network perception. Specifically, we propose novel communication scheduling protocols for model distribution and model aggregation, respectively. We have implemented TSEngine in a mainstream parameter server system and verified its effectiveness in DML-WANs testbeds. Huaman Zhou, Weibo Cai, Zonghang Li, Hong-Fang Yu, Long Luo, Gang Sun 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2020 | Online job scheduling for distributed machine learning in optical circuit switch networks
Hong-Fang Yu, Gang Sun 0001, Huaman Zhou, Zonghang Li, Shouxi Luo |
Knowl. Based Syst. | 5 |