Miao Zhang 0037

dblp:60/7041-37 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0003-1839-3079ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HyperWeave: QoS-aware GPU Overcommitment for Deep Learning Training Jobs
Yong Peng 0006, Lujia Yin, Miao Zhang 0037
IWQoS5
2026 Natural Language-Guided Autonomous Agents for Counterterrorism Simulation via Deep Reinforcement Learning
Xinmeng Li, Kai Xu 0014, Yue Hu 0016, Miao Zhang 0037, Quanjun Yin
SIMULTECH4
2026 Boosting the Performance of Decentralized Federated Learning via Catalyst Acceleration
abstract
Decentralized Federated Learning has emerged as an alternative to centralized architectures due to its faster training, privacy preservation, and reduced communication overhead. In decentralized communication, the server aggregation phase in Centralized Federated Learning shifts to the client side, which means that clients connect with each other in a peer-to-peer manner. However, compared to the centralized mode, data heterogeneity in Decentralized Federated Learning will cause larger variances between aggregated models, which leads to slow convergence in training and poor generalization performance in tests. To address these issues, we introduce Catalyst Acceleration and propose an acceleration Decentralized Federated Learning algorithm called DFedCata. It consists of two main components: the Moreau envelope function, which primarily addresses parameter inconsistencies among clients caused by data heterogeneity, and Nesterov's extrapolation step, which accelerates the aggregation phase. Theoretically, we prove the optimization error bound and generalization error bound of the algorithm, providing a further understanding of the nature of the algorithm and the theoretical perspectives on the hyperparameter choice. Empirically, we demonstrate the advantages of the proposed algorithm in both convergence speed, computational cost, and generalization performance on CIFAR10/100 and Tiny-ImageNet with various non-iid data distributions. Moreover, extensive experiments are conducted to validate the theoretical properties of DFedCata, showing strong consistency between theory and empirical observations.
Qinglun Li, Miao Zhang 0037, Yingqi Liu, Quanjun Yin, Li Shen 0008, Xiaochun Cao
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Deep Model Fusion: A Survey
abstract
Deep model fusion/merging is an emerging technique that integrates parameters or predictions from multiple deep learning (DL) models into a unified framework. It combines the abilities of different models to compensate for the biases and errors of an individual model, improving overall performance. However, deep model fusion, especially on large-scale DL models such as large language models (LLMs) and foundation models, faces several challenges, including high computational cost and interference between different heterogeneous models. In order to understand it better, we present a comprehensive survey to summarize the recent progress. We categorize existing model fusion methods as fourfold: 1) weight average (WA) averages the parameters of multiple models to obtain results closer to the optimal solution; 2) considering that direct averaging of models often yields suboptimal results, "mode connectivity" connects networks via paths of nonincreasing loss in weight spaces before the fusion. Along these paths, initial models are transformed into forms with consistent functions and better fusion effects; 3) similarly, for models with poor direct fusion results, "alignment" matches the corresponding units and merges these models, thus fully exploiting the corresponding relationships between the models; and 4) in addition to the above-mentioned methods of parameter fusion, "ensemble learning" fuses the outputs of multiple models in the inference stage to improve the accuracy and robustness of networks. In addition, we analyze the challenges of deep model fusion and illuminate the possible research directions in the future.
Yong Peng 0006, Miao Zhang 0037, Liang Ding 0006, Han Hu 0003, Li Shen 0008
IEEE Trans. Neural Networks Learn. Syst.3
2025 AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration
abstract
Multi-agent systems (MAS) based on large language models (LLMs) have demonstrated significant potential in collaborative problemsolving.However, they still face substantial challenges of low communication efficiency and suboptimal task performance, making the careful design of the agents' communication topologies particularly important.Inspired by the management theory that roles in an efficient team are often dynamically adjusted, we propose AgentDropout, which identifies redundant agents and communication across different communication rounds by optimizing the adjacency matrices of the communication graphs and eliminates them to enhance both token efficiency and task performance.Compared to state-of-the-art methods, AgentDropout achieves an average reduction of 21.6% in prompt token consumption and 18.4% in completion token consumption, along with a performance improvement of 1.14 on the tasks.Furthermore, the extended experiments demonstrate that AgentDropout achieves notable domain transferability and structure robustness, revealing its reliability and effectiveness.We release our code at https://github. com/wangzx1219/AgentDropout.
Zhexuan Wang, Xuebo Liu 0002, Liang Ding 0006, Miao Zhang 0037, Jie Liu 0001, Min Zhang 0005
ACL (1)5
2025 Integrated Data Collection and Model Retraining Optimization in UAV-Enabled ECNs
abstract
Edge-based machine learning inference typically relies on models trained on static or historical datasets, making them susceptible to performance degradation when data distribution shifts. While model retraining (continuous learning) can alleviate this issue, the integration of real-time data acquisition and model updates remains challenging due to the inherently distributed nature of data across end devices. Unmanned aerial vehicles (UAVs) offer an efficient means of gathering end-device data in edge computing networks (ECNs), thereby enabling realtime retraining. However, existing approaches design UAV-based data collection and model training in isolation, hindering realtime data-augmented training. To address this gap while account for UAVs' limited computational capacity, we propose an integrated optimization framework that coordinates data collection and model retraining through joint UAV swarm deployment and bandwidth allocation, enabling real-time model updates at the edge server. To this end, our objective is to maximize both the collected data volume and the number of training iterations, determined by data volume, transfer time, and update time, within a constrained training window. We formulate this NP-hard problem with tightly coupled variables and develop RL-EDABA, a Reinforcement Learning algorithm Embedded with Device Association and Bandwidth Allocation, enhanced by greedy strategies and convex optimization. Experiments show that RLEDABA effectively mitigates model accuracy loss with lower computational complexity compared with baseline methods.
Ding Xu 0001, Lingjie Duan, Miao Zhang 0037
ICPADS4
2025 Unveiling the Power of Multiple Gossip Steps: A Stability-Based Generalization Analysis in Decentralized Training
abstract
Decentralized training removes the centralized server, making it a communication-efficient approach that can significantly improve training efficiency, but it often suffers from degraded performance compared to centralized training. Multi-Gossip Steps (MGS) serve as a simple yet effective bridge between decentralized and centralized training, significantly reducing experiment performance gaps. However, the theoretical reasons for its effectiveness and whether this gap can be fully eliminated by MGS remain open questions. In this paper, we derive upper bounds on the generalization error and excess error of MGS using stability analysis, systematically answering these two key questions. 1). Optimization Error Reduction: MGS reduces the optimization error bound at an exponential rate, thereby exponentially tightening the generalization error bound and enabling convergence to better solutions. 2). Gap to Centralization: Even as MGS approaches infinity, a non-negligible gap in generalization error remains compared to centralized mini-batch SGD ($\mathcal{O}(T^{\frac{c\beta}{c\beta +1}}/{n m})$ in centralized and $\mathcal{O}(T^{\frac{2c\beta}{2c\beta +2}}/{n m^{\frac{1}{2c\beta +2}}})$ in decentralized). Furthermore, we provide the first unified analysis of how factors like learning rate, data heterogeneity, node count, per-node sample size, and communication topology impact the generalization of MGS under non-convex settings without the bounded gradients assumption, filling a critical theoretical gap in decentralized training. Finally, promising experiments on CIFAR datasets support our theoretical findings.
Qinglun Li, Yingqi Liu, Miao Zhang 0037, Xiaochun Cao, Quanjun Yin, Li Shen 0008
NeurIPS3
2025 Multi-agent reinforcement learning for task offloading with hybrid decision space in multi-access edge computing
Miao Zhang 0037, Quanjun Yin, Lujia Yin, Yong Peng 0006
Ad Hoc Networks2
2025 DFedGFM: Pursuing global consistency for Decentralized Federated Learning via global flatness and global momentum
Qinglun Li, Miao Zhang 0037, Tao Sun 0005, Quanjun Yin, Li Shen 0008
Neural Networks2
2025 Asymmetrically Decentralized Federated Learning
abstract
To address the communication burden and privacy concerns associated with the centralized server in Federated Learning (FL), Decentralized Federated Learning (DFL) has emerged, which discards the server with a peer-to-peer (P2P) communication framework, significantly expanding the application scenarios of FL. However, most existing DFL algorithms are based on symmetric topologies, such as ring and grid topology, which can easily lead to deadlocks and are susceptible to the impact of network link quality in practice. To address these issues, we propose DFedSGPSM, a transitional framework that converts symmetric DFL optimizers into asymmetric variants. By adopting the Push-Sum protocol in asymmetric network topologies, our framework successfully circumvents the deadlock and link-quality issues prevalent in symmetric configurations. To further validate the effectiveness of our algorithm framework, we integrate the local momentum (in DFedAvgM) and SAM (in DFedSAM) from existing symmetric DFL optimizer into DFedSGPSM to accelerate training and pursue smooth local minimum, which enables existing symmetric DFL optimizers to be seamlessly integrated into asymmetric DFL. Theoretical analysis proves that DFedSGPSM achieves a linear speedup rate of$\mathcal{O}(\frac{1}{\sqrt{nT}})$in the non-convex setting. This analysis also reveals crucial issues such as tighter upper bounds achieved with improved topological connectivity. Empirically, extensive experiments conducted on the MNIST, CIFAR10&100 datasets demonstrate the superior performance of our proposed algorithm compared to several existing SOTA optimizers in terms of generalization.
Qinglun Li, Miao Zhang 0037, Quanjun Yin, Li Shen 0008, Xiaochun Cao
IEEE Trans. Computers2
2025 The Risk of Federated Learning to Skew Fine-Tuning Features and Underperform Robustness
abstract
To tackle the scarcity and privacy issues associated with domain-specific datasets, the integration of federated learning in conjunction with fine-tuning (FT) has emerged as a practical solution. However, our findings reveal that federated learning has the risk of skewing FT features and compromising the out-of-distribution (OOD) robustness of pretrained models. By introducing three robustness indicators and conducting experiments across diverse robust datasets, we elucidate these phenomena by scrutinizing the ability of data representations, transferability, and deviations within the model. To mitigate the negative impact of practical federated learning on model robustness, we introduce a general noisy projection (GNP)-based robust algorithm, ensuring no deterioration of accuracy on the target distribution. Specifically, the key strategy for enhancing model robustness entails the transfer of robustness from the pretrained model to the fine-tuned model, coupled with adding a small amount of Gaussian noise to augment the representative capacity of the model. The comprehensive experimental results demonstrate that our approach markedly enhances the robustness across diverse scenarios, encompassing various parameter-efficient FT (PEFT) methods and confronting different levels of label distribution skew and quantity distribution skew.
Mengyao Du, Miao Zhang 0037, Yuwen Pu, Qingming Li, Shouling Ji, Quanjun Yin
IEEE Trans. Neural Networks Learn. Syst.2
2024 Low-Cost, High-Reliability Deployment for Cloud Applications With Low-Frequency Periodic Requests
abstract
Low-frequency periodic requests are common in cloud-based enterprise applications. These infrequent requests often leave microservices idle for extended periods, leading to low resource utilization. Furthermore, the randomness of response times may decrease the reliability of the cloud platform. Intuitively, the periodic nature of requests allows for the agile deployment of microservices to promptly free up occupied computing resources. Thus, the key lies in designing low-cost, high-reliability microservice deployment schemes. Traditional approaches relying on specialized expertise are impractical because of intricate interdependencies within microservice frameworks. To address this, the Microservice Deployment Problem for Low-frequency Periodic Requests (MDP-LPR) is formulated, and a Mixed Integer Programming (MIP) model is developed. A deployment framework leveraging statistical analysis and Monte Carlo simulation is proposed to ensure high reliability. Furthermore, a two-stage heuristic algorithm named Relaxation and Precision Mixed Algorithm (RPMA) is introduced to generate low-cost deployment schemes. Finally, experiments are conducted on real-world workflows. The results show that the RPMA outperforms its counterparts in generating low-cost deployment schemes, and the proposed deployment framework enables the automatic acquisition of low-cost, high-reliability deployment schemes.
Zhu Xiang, Lujia Yin, Miao Zhang 0037, Quanjun Yin
IEEE Trans. Serv. Comput.4
2023 Credit-based Differential Privacy Stochastic Model Aggregation Algorithm for Robust Federated Learning via Blockchain
abstract
By encapsulating model parameters in blocks when training machine learning models collaboratively, blockchain is recognized as a promising enabling technology to facilitate reliable federated learning under a distributed and untrusted environment. However, storing updated models of each worker in blockchain induces potential privacy risks, such as membership inference attacks. Besides, the volatile network conditions in the distributed environment may cause the deterioration of system robustness. This paper aims at addressing the privacy and robustness issues mentioned above. Specifically, a Credit-based Differential Privacy stochastic model aggregation algorithm combined with SIGN operation (Cre-DPSIGN) is adopted in our peer-to-peer network, which can realize the trade-off between privacy and accuracy. Furthermore, leveraging the transparency and tamper-proofing of blockchain, we design practical and reliable smart contracts for unbiased sampling based on the credit of workers to improve system robustness against Byzantine workers. In addition, we have demonstrated that the use of biased differential privacy mechanisms can lead to performance degradation. Therefore, we have introduced two unbiased differential privacy mechanisms and have proven their convergence and privacy guarantee. Extensive experiments conducted on MNIST datasets show that our algorithm can achieve byzantine fault tolerance rate with a private loss ϵ = 0.4. Compared with the state-of-the-art, aka DP-RSA (IJCAI-22), Cre-DPSIGN shows lower privacy loss consumption and better system robustness.
Mengyao Du, Miao Zhang 0037, Lin Liu 0018, Kai Xu 0014, Quanjun Yin
ICPP2
2022 FedHiSyn: A Hierarchical Synchronous Federated Learning Framework for Resource and Data Heterogeneity
abstract
Federated Learning (FL) enables training a global model without sharing the decentralized raw data stored on multiple devices to protect data privacy. Due to the diverse capacity of the devices, FL frameworks struggle to tackle the problems of straggler effects and outdated models. In addition, the data heterogeneity incurs severe accuracy degradation of the global model in the FL training process. To address aforementioned issues, we propose a hierarchical synchronous FL framework, i.e., FedHiSyn. FedHiSyn first clusters all available devices into a small number of categories based on their computing capacity. After a certain interval of local training, the models trained in different categories are simultaneously uploaded to a central server. Within a single category, the devices communicate the local updated model weights to each other based on a ring topology. As the efficiency of training in the ring topology prefers devices with homogeneous resources, the classification based on the computing capacity mitigates the impact of straggler effects. Besides, the combination of the synchronous update of multiple categories and the device communication within a single category help address the data heterogeneity issue while achieving high accuracy. We evaluate the proposed framework based on MNIST, EMNIST, CIFAR10 and CIFAR100 datasets and diverse heterogeneous settings of devices. Experimental results show that FedHiSyn outperforms six baseline methods, e.g., FedAvg, SCAFFOLD, and FedAT, in terms of training accuracy and efficiency.
Yue Hu 0016, Miao Zhang 0037, Ji Liu 0003, Quanjun Yin, Yong Peng 0006, Dejing Dou
ICPP3
2022 FedGosp: A Novel Framework of Gossip Federated Learning for Data Heterogeneity
abstract
Federated learning (FL) provides the possibility to solve the problem of data privacy, but it suffers much from the data heterogeneity among different participants. Currently, some promising FL algorithms improve the effectiveness of learning under the non independent-and-identically-distributed (Non-IID) data settings. However, they require a large number of communication rounds between the server and clients for an acceptable accuracy. Inspired by the training paradigm of gossip learning, this paper proposes a new FL framework, named FedGosp. It first classifies the clients into different categories based on the model weights trained by the locally stored data. Then FedGosp utilizes the communication not only between clients and the server, but also between different classes of clients themselves. This training process enables instilling knowledge about various data distributions in the passed models. We evaluate the performance of FedGosp in multiple Non-IID settings on CIFAR10 and MNIST datasets, and compare it with the recently popular algorithms such as SCAFFOLD, FedAvg and FedProx. The experimental results show that FedGosp can improve the model accuracy by 6.53% and save 5.6 × communication costs at best compared to the second-ranked baseline.
Yue Hu 0016, Miao Zhang 0037, Li Li 0064, Tao Chang, Quanjun Yin
SMC3
2022 Efficient Flow-Based Scheduling for Geo-Distributed Simulation Tasks in Collaborative Edge and Cloud Environments
abstract
Edge computing is a good complement to cloud computing for deploying large-scale geo-distributed simulation applications, which are very sensitive to the communication delay among different simulation components (also called tasks in this paper) and users. We mainly focus on the efficient scheduling of simulation components in collaborative edge and cloud environments. As components should be deployed jointly with the consideration of capacity constraints of hosts, it is actually an NP-complete multi-dimensional bin packing problem. Meanwhile, dynamic changes of component and host states require the low deployment latency of scheduling algorithms. Unfortunately, most of the existing schedulers for modern clusters are queue-based, in which tasks are scheduled sequentially, thus lacking the ability to process tightly coupled tasks jointly. Other batching-based placement algorithms are usually time-consuming. This paper describes Pond, a novel flow-based scheduler with the awareness of interactions among tasks and users as well as heterogeneous multi-dimensional resources. First, characteristics of distributed simulation tasks are analysed and the scheduling problem is formulated as a min-cost max-flow (MCMF) problem over the flow network by mapping the communication overhead among tasks and users to the costs of arcs in the network. Considering the inherent defects of existing flow-based schedulers in dealing with multi-dimensional resources, a new method based on dominant resource is proposed and some problem specific heuristics are also designed. Extensive simulation experiments based on Alibaba production trace and some random synthetic parameters are conducted. Results show that Pond can reduce the average communication cost for each task significantly in a quite low deployment latency compared with some baselines.
Miao Zhang 0037, Yong Peng 0006, Jiancheng Zhu, Quanjun Yin
IEEE Trans. Parallel Distributed Syst.1
2021 A discrete PSO-based static load balancing algorithm for distributed simulations in a cloud environment
Miao Zhang 0037, Yong Peng 0006, Quanjun Yin, Xu Xie 0005
Future Gener. Comput. Syst.1