Huaman Zhou

dblp:269/3524 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-6421-664XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Cube-fx: Mapping Taylor Expansion Onto Matrix Multiplier-Accumulators of Huawei Ascend AI Processors
abstract
Taylor expansion, a mature method for function evaluations used in Artificial Intelligence (AI) applications, approximates functions with polynomials. In addition to the function evaluations, AI applications require massive matrix multiplications, inspiring manufacturers to propose AI processors with matrix multiplier-accumulators (MACs). However, compared with the powerful Matrix MACs, the vectorized units of the AI processors cannot efficiently carry the existing Taylor expansion implementation of Single Instruction Multiple Data (SIMD) parallelism. Leveraging the Matrix MACs for Taylor expansion becomes an ideal direction. In previous studies, migrating optimized algorithms to the Matrix MACs requires matrix generation during the runtime. The generation is expensive and even cancels the accelerations brought by the Matrix MACs on the AI processors, which Taylor expansion also suffers. This article presents Cube-fx, a mapping algorithm of Taylor expansion for multiple functions onto Matrix MACs. Cube-fx expresses the building and computation in matrix multiplications without inefficient dynamic matrix generation. On Huawei Ascend processors, Cube-fx averagely achieves 1.64× speedups compared with vectorized Horner's Method with 56.38$\%$vectorized operations reduced.
Yifeng Tang, Huaman Zhou, Zhuoran Ji, Cho-Li Wang
IEEE Trans. Parallel Distributed Syst.2
2023 FedGSync: Jointly Optimized Weak Synchronization and Gradient Transmission for Fast Distributed Machine Learning in Heterogeneous WAN
abstract
Due to privacy and cost reasons, distributed machine learning in Wide-Area Networks(DML-WAN) is becoming an emerging and popular collaborative learning paradigm. However, heterogeneity in computing power and data distribution among workers in different locations has a dramatic impact on training performance, including convergence speed and learning accuracy. Most of the existing works on distributed training mechanisms either focus on computing heterogeneity or data heterogeneity, and none of them can handle both well. In this paper, we propose FedGSync, a novel distributed training mechanism to improve the training performance for DML-WAN, where computing heterogeneity and data heterogeneity usually coexist. To speed up training and improve model accuracy, FedGSync clusters workers into groups according to the similarity of their data distribution and introduce group-based weak synchronization to minimize the synchronization delays waiting for slow workers and the accuracy loss by balancing the contributions of all data distributions. To preserve data privacy and improve efficiency, FedGSync only groups workers based on principal components of gradients and design an approximate grouping mechanism based on Kmeans. To further reduce synchronization time, FedGSync prioritizes packets and uses differential transmission for gradient packets between groups. Evaluation results demonstrate that FedGSync improves convergence speed and learning accuracy under the coexistence of computing heterogeneity and data heterogeneity compared with state-of-the-art distributed training mechanisms.
Huaman Zhou, Yihong He, Long Luo, Hong-Fang Yu, Gang Sun 0001
SMC1
2023 NBSync: Parallelism of Local Computing and Global Synchronization for Fast Distributed Machine Learning in WANs
abstract
Recently, due to privacy concerns, distributed machine learning in Wide-Area Networks (DML-WANs) attracts increasing attention and has been widely deployed to promote the widespread application of intelligence services that rely on geographically distributed data. DML-WANs is essentially performing collaboratively federated learning over a combination of servers at both edge and cloud on a large spatial scale. However, efficient model training is challenging for DML-WANs because it is blocked by the high overhead of model parameter synchronization between computing servers over WANs. The reason is that there has a sequential dependency between local model computing and global model synchronization of traditional DML-WANs training methods intrinsically producing a sequential blockage between them, e.g., FedAvg. When the computing heterogeneity and the low WAN bandwidth coexist, a long block of global model synchronization prolongs the training time and leads to low utilization of local computing. Despite many efforts on alleviating synchronization overhead with novel communication technologies and synchronization methods, they still use traditional training patterns with sequential dependency and thereby have very limited improvements, such as FedAsync and ESync. In this article, we propose NBSync, a novel training algorithm for DML-WANs, which greatly speeds up the model training by the parallelism of local computing and global synchronization. NBSync employs a well-designed pipelining scheme, which can properly relax the sequential dependency of local computing and global synchronization and process them in parallel so as to overlap their operating overhead in the time dimension. NBSync also realizes flexible, differentiated and dynamical local computing for workers to maximize the overlap ratio in dynamically heterogeneous training environments. Convergence analysis shows that the convergence rate of NBSync training process is asymptotically equal to that of SSGD, and NBSync has a better convergence efficiency. We implemented the prototype of NBSync based on a popular parameter server system, i.e., MXNET's PS-LITE library, and evaluate its performance on a DML-WANs testbed. Experimental results show that NBSync speeds up training about 1.43×–2.79× than state-of-the-art distributed training algorithms (DTAs) in DML-WANs scenarios where computing heterogeneity and low WAN bandwidth coexist.
Huaman Zhou, Zonghang Li, Hong-Fang Yu, Long Luo, Gang Sun 0001
IEEE Trans. Serv. Comput.1
2022 ESync: Accelerating Intra-Domain Federated Learning in Heterogeneous Data Centers
abstract
Federated Learning (FL) serves privacy-preserving collaborative learning among multiple isolated parties, while retaining their privacy data locally. Cross-device and cross-silo FL have achieved great success in cross-domain applications, in which the scarce communication resource is the primary bottleneck. Driven by the need to combine heterogeneous machines from different parties to build a shared data center, we foundintra-domain FL, a new type of FL in which isolated parties collaborate in the shared data center, and strong computational heterogeneity becomes the primary bottleneck. To mitigate the training inefficiency caused by stragglers, this article proposes an efficient synchronization algorithmESync, which allows parties to train different iterations locally under the coordination of a novel schedulerState Server. We give the boundaries of weight divergence and optimality gap ofESync, and analyze the trade-off between convergence accuracy and communication efficiency. Extensive experiments are conducted to compareESyncwith SSGD, ASGD, DC-ASGD, FedAvg, FedAsync, TiFL, and FedDrop under strong computational heterogeneity. Numerical results show thatESyncachieves great speed up without loss of accuracy, and therefore demonstrate the effectiveness ofESyncin both training efficiency and converged accuracy.
Zonghang Li, Huaman Zhou, Tianyao Zhou, Hong-Fang Yu, Zenglin Xu, Gang Sun 0001
IEEE Trans. Serv. Comput.2
2021 DGT: A contribution-aware differential gradient transmission mechanism for distributed machine learning
Huaman Zhou, Zonghang Li, Qingqing Cai, Hong-Fang Yu, Shouxi Luo, Long Luo, Gang Sun 0001
Future Gener. Comput. Syst.1
2021 TSEngine: Enable Efficient Communication Overlay in Distributed Machine Learning in WANs
abstract
In recent years, distributed machine learning in WANs (DML-WANs), i.e., collaboratively training a high-quality ML model cross geo-distributed micro-clouds or edge devices, has attracted attention and been widely applied. Compared with cloud-centric training, DML-WANs avoids the high cost of transferring large amounts of raw data to a central cloud and privacy concerns. However, performing DML-WANs still faces challenges. Model synchronization, an essential step of DML-WANs, is accompanied by a lot of model communication cross limited-bandwidth WANs, which generates high communication overhead. Moreover, the parameter server system, which has been widely used, performs model synchronization in a centralized manner, resulting in serious communication in-cast problem. Such communication in-cast further raises the communication overhead, leading to the low efficiency of DML-WANs. To alleviate the communication in-cast, existing researches attempt to build tree-based communication overlays over the parameter server and workers. However, we identify that these approaches can not adapt to the dynamic and heterogeneous network of DML-WANs, resulting in insufficient improvements. This paper proposes TSEngine, an adaptive communication scheduler for efficient communication overlay of the parameter server system in DML-WANs. Its core idea is to dynamically schedule the communication logic over the parameter server and workers based on the active network perception. Specifically, we propose novel communication scheduling protocols for model distribution and model aggregation, respectively. We have implemented TSEngine in a mainstream parameter server system and verified its effectiveness in DML-WANs testbeds.
Huaman Zhou, Weibo Cai, Zonghang Li, Hong-Fang Yu, Long Luo, Gang Sun 0001
IEEE Trans. Netw. Serv. Manag.1
2020 Online job scheduling for distributed machine learning in optical circuit switch networks
Hong-Fang Yu, Gang Sun 0001, Huaman Zhou, Zonghang Li, Shouxi Luo
Knowl. Based Syst.4