Benteng Zhang

dblp:360/2703 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient and Robust Federated Learning via Synergistic Aggregation on Heterogeneous Devices
Yihan Chen 0002, Yingchi Mao, Benteng Zhang, Xiaoming He 0004, Miao Du, Jie Wu 0001
ICC4
2026 Mitigating grid fluctuations in renewable power systems via vehicle-to-grid: A hybrid meta-heuristic optimization framework
Benteng Zhang, Yuanjun Guo, Kang Li 0002, Zhile Yang
Expert Syst. Appl.2
2025 Energy-Efficient Federated Learning via Dynamic Distillation and Cloud-Network Collaboration
abstract
The extensive local training in Hierarchical Federated Learning (HFL) imposes a substantial computational energy burden on end devices, a problem intensified by inherent system and data heterogeneity. While prior works attempt to mitigate this by using heterogeneous models or adjusting local training, they often suffer from critical drawbacks such as accuracy degradation from update biases and an inability to adapt to the dynamic nature of device resources. This paper introduces FedE2AD (Federated Energy-Efficient Adaptive Distillation), a novel cloud-network collaboration framework that leverages dynamic distillation to holistically optimize the energy-accuracy balance. At its core, FedE2AD implements this collaboration through a multi-level optimization approach. At the cloud layer, a Dynamic Model Allocation strategy intelligently assigns model architectures by assessing device status from static, dynamic, and data-centric perspectives. At the device layer, Variable Local Iterations enable real-time adaptation to fluctuating computational power. Crucially, to counteract model divergence, FedE2AD employs a Dual Knowledge Sharing mechanism at the edge layer, which uniquely combines direct aggregation of shared structures with data-free model distillation to ensure robust knowledge transfer. Experiments conducted on simulation platforms show that FedE2AD markedly outperforms existing methods. For instance, on the CIFAR-10 dataset under strong heterogeneity, it reduces single-round computation energy by 21.1% and increases final model accuracy by 1.42% compared to HDHRFL.
Yihan Chen 0002, Benteng Zhang, Xiaoming He 0004, Miao Du, Yingchi Mao
ICNP4
2025 LLM-Driven Cloud-Edge Collaboration for Resilient Multi-UAV Task Planning
abstract
Large Language Model (LLM)-driven multi-agent task planning offers a promising approach for automating complex missions, particularly in critical domains like disaster response. Nevertheless, the efficacy of existing planners is often compromised in dynamic environments. This vulnerability primarily stems from their inability to enforce complex, non-linear task dependencies, coupled with a lack of mechanisms to recover from execution failures. To address these gaps in reliability and resilience, this paper introduces ReFlex-LLM, a novel cloud-edge framework for autonomic multi-UAV task planning. ReFlex-LLM leverages a Directed Acyclic Graph (DAG) to explicitly model and guarantee the logical correctness of task sequencing, thereby preventing foundational planning errors. In parallel, ReFlex-LLM incorporates a closed-loop feedback mechanism, allowing the central planner to receive execution status from edge UAVs and to adapt the mission plan in response to unforeseen contingencies. Extensive simulations on complex, dependency-heavy tasks demonstrate that ReFlex-LLM significantly outperforms state-of-the-art baselines, boosting the overall mission success rate by over 20%.
Xuan Ling, Yingchi Mao, Benteng Zhang, Xiaoming He 0004
ICPADS4
2025 HFHEMS: Energy-Efficient Hierarchical Federated Learning via Model Distillation
abstract
Extensive local training in Hierarchical Federated Learning (HFL) imposes high energy demands on resourceconstrained devices, a problem exacerbated by system heterogeneity which also causes performance variability. To address this, we propose HFHEMS, a novel method utilizing heterogeneous models and distillation to improve energy efficiency in HFL. HFHEMS employs dynamic model allocation to tailor computational loads to device capabilities and uses variable local iterations for real-time training adjustment. To preserve model accuracy, it integrates a dual knowledge sharing strategy with data-free distillation, enhancing knowledge transfer. Experiments confirm HFHEMS significantly reduces computation energy while maintaining robust accuracy, thus achieving a superior energy-accuracy trade-off.
Yihan Chen 0002, Yingchi Mao, Benteng Zhang, Qinxiao Deng, Xiang Li 0209, Xiaoming He 0004
IWQoS3
2025 Efficient Zero-Cost Neural Architecture Search for Personalized AI Systems in Cloud-Edge Networks
abstract
Neural Architecture Search (NAS) can discover the optimal neural network architecture within a given SuperNet through automated search, which can improve model performance and reduce computational overhead on resource-constrained devices. Due to the vast SuperNet requiring substantial computational resources for training and evaluation, the search process is costly and difficult to apply directly on End Devices (EDs) with limited computational resources. Moreover, existing methods utilize zero-cost proxies to reduce computational costs in NAS, but overlook limited computational resources on EDs and waste a large amount of computational resources on the cloud server. Deploying NAS on the cloud server can effectively address this issue. The cloud server is used to search for the optimal Subnet, and EDs only need to train the Subnet based on local data. To this end, we propose a nonlinear aggregation-based Neural Architecture Search method based on Feature and Gradient zero-cost proxies (FG-NAS). Specifically, EDs upload local data characteristics to the cloud server, and then the cloud server uses FG-NAS to obtain an optimal SubNet model from the SuperNet based on the uploaded data characteristics. Finally, the cloud server sends the optimal SubNet to EDs, which can reduce the computational burden on EDs. Furthermore, FG-NAS evaluates the accuracy of neural architectures by considering both feature proxies during forward propagation and gradient proxies during backward propagation. Experiments on three datasets demonstrate that compared to current mainstream zero-cost proxy methods, FG-NAS can improve evaluation accuracy by an average of 1.04% and reduce single-network evaluation time by up to 2.45%.
Yingchi Mao, Benteng Zhang, Yihan Chen 0002, Yuchu Chen, Jie Wu 0001
MASS3
2025 Learnable Cloud-Guided LLM Quantization for Resource-Constrained Edge Devices
Qinxiao Deng, Tianfu Pang, Benteng Zhang, Bingbing Nie, Xiaoming He 0004, Yingchi Mao, Jie Wu 0001
NPC (1)3
2025 Temperature-Aware Adaptive Federated Distillation for Energy-Constrained AIoT with Non-IID Data
abstract
Federated Learning (FL) can help multiple Internet of Things (IoT) devices to collaboratively train a machine learning model to provide intelligent services and applications (FL-AIoT). Due to IoT devices' limited storage capacity and energy, fresh data collected by devices often overwrites outdated data and establishes a heterogeneous data distribution. This causes the global model to forget outdated data's characteristics (i.e., catastrophic forgetting). Existing methods incorporate knowledge distillation into FL (i.e., federated distillation, FD) to extract and integrate characteristics from both fresh and outdated data, but they use fixed distillation temperatures for different devices, which overlooks that fixed distillation temperatures cannot match the heterogeneous data distribution on different devices and degrades global model accuracy. To this end, we propose a Federated Dynamic Decoupled Distillation method based on Logits distribution (Fed3DL). Specifically, Fed3DL utilizes decoupled distillation to mitigate catastrophic forgetting. To alleviate the impact of heterogeneous data distributions, Fed3DL novelly builds an adaptive temperature-aware mechanism to dynamically adjust the distillation temperature of each device based on the distribution of Logits. Additionally, Fed3DL introduces a regularization term into the local distillation loss to reduce inter-class characteristics disparity and improve model accuracy. Experiments on two datasets show that compared with the best of the 5 baselines, Fed3DL can improve the global model accuracy by an average of 3.40 %, reduce the forgetting rate by an average of 4.98 %, and achieve the lowest inter-class accuracy disparity.
Yingchi Mao, Jiakai Zhang, Litao Qu, Benteng Zhang, Xiaoming He 0004
VTC2025-Spring5
2025 Overcoming Forgetting Using Adaptive Federated Learning for IIoT Devices With Non-IID Data
abstract
In real-world Industrial Internet of Things (IIoT) scenarios, due to the limited storage capacity of IIoT devices, fresh data continuously received by diverse devices will overwrite the outdated data and change the local data distribution. However, state-of-the-art studies have demonstrated that federated learning tends to focus on training with fresh data, and the latest global model may forget the historical update directions (i.e., catastrophic forgetting). This issue can significantly degrade the global model accuracy. Existing methods primarily focus on integrating outdated data characteristics into fresh data but overlook the large parameter update gap between global and local models during global aggregation. This gap can cause the global model updates to deviate from the optimal direction. To this end, we propose a federated adaptive weighted aggregation method based on model consistency (FedAWAC). Specifically, FedAWAC measures the model consistency on devices and dynamically adjusts the aggregation weights of each local model, thereby guiding the global model toward optimal updates. Furthermore, FedAWAC integrates$\mathcal {M}$historical global models most correlated to the latest global model on the cloud server to overcome catastrophic forgetting. Experiments on four different datasets (nonidentically and independently distributed settings) indicate that compared to five baselines, FedAWAC can improve global model accuracy by an average of 1.86%, reduce the forgetting rate by an average of 3.91%, and save average memory usage by up to 2.57 GB.
Benteng Zhang, Yingchi Mao, Yihan Chen 0002, Tasiu Muazu, Xiaoming He 0004, Jie Wu 0001
IEEE Internet Things J.1
2025 Adaptive Knowledge Transfer for Federated Learning With Large Models on Edge IoT Devices
abstract
Cloud-centric deployment of large models brings privacy concerns and communication burdens to IoT devices. Federated Learning (FL) can help numerous IoT devices collaboratively train a large model in a distributed manner to provide intelligent services and applications (AIoT). However, due to the limited storage capacity of IoT devices, fresh data collected by IoT devices often overwrites outdated data and establishes Non-IID data distributions. Moreover, state-of-the-art studies have indicated that FL tends to use fresh data for model training. This causes the global model to forget outdated data’s characteristics (i.e., catastrophic forgetting). Some methods utilize Knowledge Distillation (KD) to extract and integrate characteristics from both fresh data and outdated data. However, existing KD-based methods use fixed distillation temperatures for different IoT devices, which overlooks that fixed distillation temperatures cannot match the data distributions on different IoT devices and may degrade global model accuracy. To this end, we propose a Federated Dynamic Decoupled Distillation method based on Logits distribution (Fed3DL). Specifically, Fed3DL utilizes decoupled distillation to mitigate catastrophic forgetting. To dynamically adjust the distillation temperature of each IoT device, Fed3DL novelly builds an adaptive temperature-aware mechanism based on the Logits distribution. Furthermore, Fed3DL introduces a regularization term into the local distillation loss to improve global model accuracy. Experiments on four datasets show that Fed3DL can improve the global model accuracy by an average of 3.19%, reduce the forgetting rate by an average of 4.46%, and achieve the lowest inter-class accuracy disparity.
Benteng Zhang, Yingchi Mao, Xiang Li 0209, Xiaoming He 0004, Jie Wu 0001
IEEE Internet Things J.1
2025 Balancing Privacy and Accuracy Using Significant Gradient Protection in Federated Learning
abstract
Previous state-of-the-art studies have demonstrated that adversaries can access sensitive user data by membership inference attacks (MIAs) in Federated Learning (FL). Introducing differential privacy (DP) into the FL framework is an effective way to enhance the privacy of FL. Nevertheless, in differentially private federated learning (DP-FL), local gradients become excessively sparse in certain training rounds. Especially when training with low privacy budgets, there is a risk of introducing excessive noise into clients’ gradients. This issue can lead to a significant degradation in the accuracy of the global model. Thus, how to balance the user's privacy and global model accuracy becomes a challenge in DP-FL. To this end, we propose an approach, known as differential privacy federated aggregation, based on significant gradient protection (DP-FedASGP). DP-FedASGP can mitigate excessive noises by protecting significant gradients and accelerate the convergence of the global model by calculating dynamic aggregation weights for gradients. Experimental results show that DP-FedASGP achieves comparable privacy protection effects to DP-FedAvg and cpSGD (communication-private SGD based on gradient quantization) but outperforms DP-FedSNLC (sparse noise based on clipping losses and privacy budget costs) and FedSMP (sparsified model perturbation). Furthermore, the average global test accuracy of DP-FedASGP across four datasets and three models is about$2.62$%,$4.71$%,$0.45$%, and$0.19$% higher than the above methods, respectively. These improvements indicate that DP-FedASGP is a promising approach for balancing the privacy and accuracy of DP-FL.
Benteng Zhang, Yingchi Mao, Xiaoming He 0004, Huawei Huang, Jie Wu 0001
IEEE Trans. Computers1
2024 FedMHC: Overcoming Dimensionality and Communication Challenges for Personalized Federated Learning Using Model Head Clustering
abstract
In real Internet of Things (IoT) environments, IoT devices vary widely in data types and needs. IoT devices participating in Personalized Federated Learning (PFL) all have their own unique data characteristics, but there exist some similarities. Previous work enhances Personalized Local Models (PLMs)’ performance by clustering IoT devices’ PLM while ignoring dimensionality and communication volume, resulting in lower Global Model (GM) accuracy and PLMs’ performance. To this end, we propose a Personalized Federated Learning method based on Model Head Clustering (FedMHC). Specifically, FedMHC groups IoT devices with similar data characteristics and distributes different GMs to different device groups. FedMHC allows each IoT device to obtain a GM that best fits its local data characteristics and guides PLM training. FedMHC only clusters model head parameters on the server side. Thus, the edge server only needs to transmit the head parameters and a single shared feature extractor parameters during communication with IoT devices. The improvement can effectively address the issues of dimensionality and high communication volume. Experiments on CIFAR-100, Tiny-ImageNet, and AG News datasets demonstrate that FedMHC enhances the model accuracy by 1.79% and 5.9% in pathological heterogeneous scenarios, and by 1.43%, 0.92%, and 0.94% in practical heterogeneous scenarios, compared to the topperforming methods among 9 baselines.
Yingchi Mao, Xiaoming He 0004, Benteng Zhang, Feng Mao, Jie Wu 0001
HPCC5
2023 Optimizing Privacy-Accuracy Trade-off in DP-FL via Significant Gradient Perturbation
abstract
In federated learning with differential privacy, an obvious phenomenon of local gradient sparsification emerges in some training rounds. When training with low privacy budgets, there is a risk of excessive noise being introduced into the uploaded gradients, leading to a significant decrease in the accuracy of the global model. To tackle the trade-off between privacy protection and model accuracy with low privacy budgets, we propose a differential privacy federated aggregation method based on gradient sparsification (DP-FedAGS), which not only prevents excessive noise addition by protecting only significant gradients, but also accelerates global model convergence by dynamically calculating the weight of the gradient. Experimental results indicate that DP-FedAGS achieves comparable privacy protection to DP-FedAvg and cpSGD, while outperforming DPFedSNLC. Moreover, our approach respectively attains an approximate average test accuracy improvement of $2 .45 \%, 4 . 79$% and $0 . 29$% over the above three methods, rendering DP-FedAGS a promising approach for exploring a balance between privacy protection and model accuracy.
Benteng Zhang, Yingchi Mao, Zijian Tu, Xiaoming He 0004, Ping Ping, Jie Wu 0001
MSN1
2023 Green Resource Allocation with DDPG for Knowledge Learning in Digital Twin-enabled Edges
abstract
In the era of Information and Communication, big data is rapidly generated due to the increasing data-driven applications in Internet of Things (IoT). Effectively processing such data, e.g., knowledge learning, on resource-limited IoT becomes a challenge. In this paper, we introduce a digital twin-enabled IoT, in order to achieve hyper-connected experience, green communication, and sustainable computing. Although knowledge learning benefits from the proposed system, system latency and energy consumption are still our focus in the distributed learning architecture. To this end, we leverage Deep Reinforcement Learning (DRL) to present the deep deterministic policy gradient with double actors and double critics (D4PG) to manage the multi-dimensional resources, i.e., CPU cycles, DT models, and communication bandwidths, enhancing the exploration ability and improving the inaccurate value estimation of agents in continuous action spaces. Extensive experimental results prove that the proposed architecture can efficiently conduct knowledge learning, and our intelligent scheme can effectively improve the system efficiency.
Xiaoming He 0004, Yingchi Mao, Yinqiu Liu, Benteng Zhang, Yan Hong 0002
VTC Fall4