Dewen Qiao

dblp:278/0807 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
19since 2021 · last 2027
0000-0002-6909-838XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Computer networks · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2027 PURE: Boundary-aware pruning with reliable statistical enhancement for dataset distillation
Ruxue Bai, Minyu Liu, Ruihong Xiu, Dewen Qiao, Junqing Le, Xiaofeng Liao 0001
Expert Syst. Appl.4
2026 EECF: An edge-end collaborative framework with optimized lightweight model
Dewen Qiao, Zhenyan Wang, Jiamiao Liu, Xuetao Chen, Di Zhang 0011, Maolan Zhang
Expert Syst. Appl.1
2026 FedFA: Efficient federated large language models with feature adapters
Lijia Zhang, Songtao Guo, Dewen Qiao
Expert Syst. Appl.4
2025 FedSODA: Federated Fine-Tuning of LLMs via Similarity Group Pruning and Orchestrated Distillation Alignment
abstract
Federated fine-tuning (FFT) of large language models (LLMs) has recently emerged as a promising solution to enable domain-specific adaptation while preserving data privacy. Despite its benefits, FFT on resource-constrained clients relies on the high computational and memory demands of full-model fine-tuning, which limits the potential advancement. This paper presents FedSODA, a resource-efficient FFT framework that enables clients to adapt LLMs without accessing or storing the full model. Specifically, we first propose a similarity group pruning (SGP) module, which prunes redundant layers from the full LLM while retaining the most critical layers to preserve the model performance. Moreover, we introduce an orchestrated distillation alignment (ODA) module to reduce gradient divergence between the sub-LLM and the full LLM during FFT. Through the use of the QLoRA, clients only need to deploy quantized sub-LLMs and fine-tune lightweight adapters, significantly reducing local resource requirements. We conduct extensive experiments on three open-source LLMs across a variety of downstream tasks. The experimental results demonstrate that FedSODA reduces communication overhead by an average of 70.6%, decreases storage usage by 75.6%, and improves task accuracy by 3.1%, making it highly suitable for practical FFT applications under resource constraints.
Manning Zhu, Songtao Guo, Pengzhan Zhou, Yansong Ning, Chang Han, Dewen Qiao
ECAI6
2025 Partitioned Collaborative Inference for On-Device Models via Evolutionary Reinforcement Learning
abstract
The growing demand for intelligent mobile applications has made the deployment and operation of Deep Neural Networks (DNNs) on mobile Edge Devices (EDs) increasingly essential. However, the limited computational resources of EDs often result in significant energy consumption and compromised inference quality. To address these challenges, we propose a Partitioned Collaborative Inference (PCI) system that reduces on-device model inference costs by distributing the inference process across multiple EDs and MEC servers. To dynamically model the relationships between computing nodes, inference tasks, and resources, we employ Graph Neural Networks to construct the current state representation of the system. Furthermore, we develop a Cross-Entropy Method (CEM) based Evolutionary Reinforcement Learning algorithm, which leverages negative temporal difference (TD) error as a population fitness metric to generate elite individuals. The elite produces high-quality samples to improve learning efficiency, thereby obtaining optimal partitioned collaborative inference decisions and resource allocation in highly dynamic and complex search spaces. Extensive simulations demonstrate that the proposed approach significantly outperforms existing methods and benchmark schemes, achieving a 57. 5% increase in the inference task completion rate and a 65.7% reduction in system costs.
Lin Tan 0011, Pengzhan Zhou, Songtao Guo, Jun Zhao 0007, Zhufang Kuang, Dewen Qiao, Lu Yang 0012
ICDCS6
2025 Resources-Efficient Accelerated K-Asynchronous Adaptive Federated Learning
abstract
Asynchronous Federated Learning (AFL) has garnered significant attention in edge computing (EC) due to its adaptability. However, the frequent communication overhead inherent in AFL poses a critical challenge in resource-constrained EC settings. To address this challenge, we propose REAFL, an innovative framework that integrates the strategic selection of the optimal set of K devices with momentum gradient to reduce resource costs and enhance model performance. We begin by establishing a mathematical framework that characterizes the relationship between the selection of K devices and the global loss of AFL. Leveraging this theoretical insight, we formulate an optimization problem aimed at minimizing global loss under a resource budget, with a focus on determining the optimal device subset. Recognizing the dynamic nature of device selection, we incorporate Deep Reinforcement Learning (DRL) into REAFL to effectively learn the best device selection strategy while also considering client utility. Extensive experimental evaluations demonstrate the superior performance of REAFL over existing benchmarks, achieving up to a 24.03% improvement in model accuracy, a 61.07% reduction in energy consumption, and a 61.34% decrease in communication rounds.
Maolan Zhang, Zhenyan Wang, Jiamiao Liu, Xuetao Chen, Dewen Qiao
IJCNN7
2025 EMAFL: Evolutionary Momentum Auxiliary Adaptive Accelerating Federated Learning
abstract
The utilization of federated learning (FL) has witnessed notable advancements in the domain of edge computing (EC). However, limited edge resources and heterogeneous devices restrict the accelerated training of the FL model. To address this issue, we introduce the biological evolutionary mechanism and momentum gradient descent (MGD) update approach into FL, called the EMAFL scheme, aiming to achieve accelerated model training and maximize resource utilization, simultaneously. Specifically, we first update the local model with particle swarm optimization (PSO) for each device and perform MGD on the updated local model. Next, by a toy example, we illustrate the necessity of adopting the different number of local iterations for heterogeneous devices in a resource-limited environment. Analytical convergence of the EMAFL scheme, premised on a delineated resource budget is subsequently explored. This yields a mathematical delineation correlating the quantity of local iterations for heterogeneous devices with the optimal model parameters. Predicated on the prior theoretical examinations, an adaptive control algorithm is devised to ascertain the local iteration count pertinent to each device following every communication round. Finally, through a lot of experiments compared with the benchmarks, the advantages of EMAFL in model accuracy, resource consumption, and Non-IID issues are verified.
Dewen Qiao, Songtao Guo, Xuetao Chen, Pengzhan Zhou, Di Zhang 0011
IEEE Internet Things J.1
2025 EESyn-CTP: Edge-End Collaboration for Patient-Friendly CTP Image Synthesis
abstract
In the field of medical imaging driven by the Internet of Things (IoT), with the rapid growth of the number of medical devices and the widespread application of edge computing (EC) technology, efficient collaborative computing on resource-constrained end medical devices has become the key to improving diagnostic efficiency, thereby bringing a more patient-friendly diagnosis and treatment experience. Computed tomography perfusion (CTP) images play an irreplaceable role in the assessment of brain tissue ischemia in patients with acute ischemic stroke (AIS), with high diagnostic accuracy in identifying ischemic lesions and distinguishing infarction from penumbra, but it has the disadvantages of high radiation dose and high cost. To this end, we propose a CTP image synthesis framework based on edge-end collaboration (EESyn-CTP), which aims to use non-contrast CT (NCCT), CT angiography (CTA), and delayed CTA (CTA+8s) images to synthesize CTP images with arbitrary time to optimize AIS diagnosis. The framework consists of two stages: the pre-training stage on the edge server and the fine-tuning stage on the end device. Specifically, we first deploy a temporal residual generative network, t-UNet, on the edge server for pre-training. This process utilizes multiple CTP images, which share similar perfusion features with CTA, CTA+8s, and NCCT images, to effectively learn the gap in perfusion information between the inputs and outputs. Subsequently, the pre-trained t-UNet model parameters are frozen and broadcast to the edge medical device. A UNet adapter is introduced before the model, and fine-tuning is performed on the adapter weights using real NCCT, CTA, and CTA+8s images as input. This approach facilitates the synthesis of CTP images at arbitrary time points. Finally, experiments on an internal data set showed that the quality of Syn-CTP images synthesized by the EESyn-CTP framework is comparable to that of real CTP images and significantly reduces computation latency and energy overhead.
Dewen Qiao, Songtao Guo, Yu Liu 0021, Qiaoqiao Ding, Xiaoqun Zhang, Xuetao Chen
IEEE Internet Things J.3
2025 Tri-AFLLM: Resource-Efficient Adaptive Asynchronous Accelerated Federated LLMs
abstract
The local deployment of federated large language models (FLLM) has further advanced the development of edge intelligence. However, the resource constraints of end devices, device heterogeneity, and the non-independent and identically distributed (Non-IID) nature of data pose significant challenges to the application of FLLM. To address this issue, we propose an Adaptive Asynchronous Accelerated FLLM (Tri-AFLLM) algorithm to achieve the efficient utilization of limited resources and improve model accuracy in the edge computing (EC) scenarios. Specifically, Tri-AFLLM first ships an off-the-shelf LLM, i.e., CLIP, to each end device, keeping the backbone parameters frozen and updating only the parameters of the adapter containing two linear transformation layers by using momentum gradient descent (MGD). Next, a toy example is provided to illustrate the necessity of using different numbers of local iterations for heterogeneous devices in resource-constrained environments. Subsequently, the convergence bound of the Tri-AFLLM under a given resource budget is discussed. Then, we formulated the bound into a resource consumption minimization problem with the number of local iterations as the optimization variable under a given model accuracy to mitigate the contribution disparity of local models to the global aggregation. Finally, extensive experiments are conducted to validate the superiority of Tri-AFLLM in terms of resource consumption, model accuracy, and addressing the Non-IID problem.
Dewen Qiao, Yu Liu 0021, Xuetao Chen, Fuyuan Song, Zheng Qin 0001, Wenqiang Jin
IEEE Trans. Circuits Syst. Video Technol.1
2025 ASMAFL: Adaptive Staleness-Aware Momentum Asynchronous Federated Learning in Edge Computing
abstract
Compared with synchronous federated learning (FL), asynchronous FL (AFL) has attracted more and more attention in edge computing (EC) fields because of its strong adaptability to heterogeneous application scenarios. However, the non-independent and identically distributed (Non-IID) data across devices and the staleness-aware estimation of unreliable wireless connections and limited edge resources make it much more difficult to achieve better AFL-related applications. To handle this problem, we propose anAdaptiveStaleness-awareMomentumAcceleratedAFL(ASMAFL) algorithm to reduce the resources consumption of heterogeneous wireless communication EC (WCEC) scenarios, as well as decrease the negative impact of Non-IID data for model training. Specifically, we first introduce the staleness-aware parameter and a unified momentum gradient descent (GD) framework to reformulate AFL. Then, we establish global convergence properties of AFL, derive an upper bound on AFL convergence rate, and find that the bound is related to the staleness-aware parameter and Non-IIDness. Next, we formulate the bound into a minimization problem of resource consumption under given model accuracy, and the corresponding staleness-aware parameter of devices will be recomputed after each asynchronous aggregation to eliminate the differences of local models’ contribution to global model aggregation. Finally, extensive experiments are carried out to validate the superiority of ASMAFL in model accuracy, convergence rate, resources consumption, Non-IID issue, etc.
Dewen Qiao, Songtao Guo, Jun Zhao 0007, Junqing Le, Pengzhan Zhou, Xuetao Chen
IEEE Trans. Mob. Comput.1
2025 AMFL: Resource-Efficient Adaptive Metaverse-Based Federated Learning for the Human-Centric Augmented Reality Applications
abstract
The emergence of 5G technology has enabled the development of Metaverse applications that provide users with immersive experiences through augmented reality (AR) devices, and the integration of federated learning (FL) with the Metaverse AR (MAR) systems can enable many edge intelligence services in 5G. However, the presence of nonindependent and identically distributed (Non-IID) data across all AR users' devices, coupled with limited edge communication resources, makes it challenging to achieve human-centric Metaverse-related applications such as target detection or image classification that combine virtual content with real-world. To address these challenges, we propose a novel adaptive resource-efficient Metaverse-based FL (AMFL) algorithm for AR applications that mitigates the negative effect of Non-IID data and reduces resource costs as well as improves the quality of experience (QoE). We first analyze the impact of wireless communication factors such as CPU frequency, bandwidth, and transmission power on FL training performance by a toy example in the MAR systems. Based on this analysis, furthermore, we establish a Non-IID degree, model accuracy, and resource consumption-related QoE maximization problem under given resource budgets, which is a stochastic optimization problem with strongly coupled variables, including bandwidth, CPU frequency, and transmission power. Guided by the theoretical analysis, to solve this issue, AMFL employs a deep reinforcement learning (DRL)-based method to adaptively allocate resources. Numerical results demonstrate that AMFL can significantly improve the QoE by up to 30.28%, and reduce communication round and energy costs by up to 81.08% and 72.20%, respectively, even under the worst Non-IID case, compared to benchmarks.
Dewen Qiao, Liangxin Qian, Songtao Guo, Jun Zhao 0007, Pengzhan Zhou
IEEE Trans. Neural Networks Learn. Syst.1
2024 RCFL-GAN: Resource-Constrained Federated Learning with Generative Adversarial Networks
abstract
Generative Adversarial Networks (GANs) training in federated learning (FL) is becoming increasingly popular in solving practical applications. However, the training is hampered by non-independent and identically distributed (non-IID) data and resource constraints. The issue of client-drift caused by non-IID data may lead to the model’s poor performance and failure to converge to the global optimum. To address these issues, we propose an adaptive RCFL-GAN framework that combines federated learning to train GANs and deep reinforcement learning (DRL) to adaptively select better-performing clients for aggregation and optimize client-local training epochs. The evaluation of contribution weights for each local model is achieved through the utilization of the maximum mean square deviation (MMD) score. The minimization issue of the RCFL-GAN model loss is formulated in a specified resource budget. Additionally, a theoretical analysis is conducted to examine the impact of client selection and the number of local training epochs on the training performance of the RCFL-GAN model. According to experimental results, our RCFL-GAN framework can improve learning performance while addressing the problems caused by non-IID data and limited resources in FL.
Yuyan Quan, Songtao Guo, Dewen Qiao
CSCWD3
2024 Reliability-Aware SFC Scheduling in Container Environment via Priority-Based Node Selection
abstract
Advances in containerization technology and edge device performance enable applications to run on a wide range of devices through virtualization, enhancing service quality in decentralized edge networks. However, edge devices often lack the computational power of cloud infrastructure and may experience connection fluctuations, which makes node reliability crucial when providing Virtual Network Functions (VNFs). To provide Service Function Chain (SFC) which combines a series of ordered VNFs, it is necessary to determine the redundancy of VNFs to meet reliability requirement and decide whether to deal with these VNF requests immediately or defer them. Therefore, this paper addresses this problem and formulates it as a reliability-aware SFC scheduling problem in container environment (RASCE) and prove it to be NP-hard. To solve this problem, we propose a reliability-aware scheduling algorithm via priority-based node selection (SSAP) using Deep Reinforcement Learning (DRL), which consists of long-task prioritization redundancy strategy considering dynamic node reliability, priority-based node selection, and SFC scheduling based on DRL. The simulation demonstrates that our approach can enhance the success rate by a minimum of 5.78% in comparison to the state-of-the-art algorithm.
Longzhi Dai, Songtao Guo, Guiyan Liu, Dewen Qiao
MSN4
2024 HfedPES: Hierarchical Personalized Federated Learning with Edge Selection
abstract
Federated learning may protect user privacy, reduce the transmission of a large amount of raw data, and is more compatible with smart home applications. Current federated learning faces two major problems including non-independent and identically (Non-IID) distributed data and high communication overhead. Personalized federated learning is a good method to deal with Non-IID data, but current personalized federated learning methods overlook the shared features of users' living habits in the same region. Hierarchical federated learning can reduce traffic on the core network, but its potential for personalization for smart home applications has not been considered. Therefore, to address these issues simultaneously, we propose hierarchical personalized federated learning. Specifically, we adopt a three-layer federated learning architecture of cloud-edge-client. On this basis, we use differential learning classification loss (DLCL), hierarchical balance loss (HBL) and balanced edge data selection (BEDS) methods to achieve the personalization of models on both the device side and the edge side. Finally, our experiments demonstrate that compared to state-of-the-art federated learning methods, hierarchical personalized federated learning has improvements in model accuracy and communication overhead.
Kunhong He, Pengzhan Zhou, Yijun Zhai, Yuepeng He, Lin Tan 0011, Dewen Qiao, Songtao Guo
MSN6
2024 Afl-gan: adaptive federated learning for generative adversarial network with resource constraints
Yuyan Quan, Songtao Guo, Dewen Qiao
CCF Trans. Pervasive Comput. Interact.3
2024 Parameters optimization and precision enhancement of Takagi-Sugeno fuzzy neural network
Dewen Qiao, Pengzhan Zhou, Songtao Guo
Soft Comput.1
2023 FedSG: Subgraph Federated Learning on Multiple Non-IID Graphs
abstract
Most of federated learning (FL) researches mainly focus on image and voice data at the expense of graph data. However, Graph Federated Learning (GFL) is specialized for FL on graph, and received little attention. Subgraph FL is a branch of GFL. In the Subgraph FL situation, a graph is not stored centrally but is distributed among clients as multiple subgraphs. Each client owns a subgraph of the original graph and faces a unique challenge, i.e., missing information cross clients. Ignoring the missing information cross subgraphs will result in deterioration of the performance of the local model. In this paper, we consider data heterogeneity and bring up a more practical problem, i.e., the subgraphs on the clients are from multiple Non-IID graphs rather than the same global graph. Then, to address the issues, we propose a subgraph FL framework FedSG which can learn a personalized model for each client, benefiting from its ability to effectively separate and combine topology information and feature information among the subgraphs. Finally, our experimental results show that FedSG achieves higher accuracy performance and faster convergence, compared with the existing approaches.
Yingcheng Wang, Songtao Guo, Dewen Qiao
MSN3
2022 HeteFL: Network-Aware Federated Learning Optimization in Heterogeneous MEC-Enabled Internet of Things
abstract
Federated learning (FL) is an effective paradigm for training a machine-learning model based on data distributed at a large quantity of users in Internet of Things (IoT) without sharing their raw data. However, federated optimization of the global model in heterogeneous IoT—while considering the heterogeneity among users and limited network constraints—remains to be an open challenge. In this article, we propose a novel adaptive federated optimization algorithm, Adp-FedProx, to achieve the optimal learning performance within the limited computation and communication resources at the edge. In particular, we analyze how the training loss is affected by each user’s global update frequency and the time and energy used for learning by obtaining the novel convergence bound of federated training loss in heterogeneous IoT. With our proposed algorithm, all users can dynamically adjust their number of local iterations in each global interval and will not drop out during the training process for resource exhaustion, so as to impair the negative effect of heterogeneity among users and guarantee the convergence of the training model. In addition, we can get the optimal learning performance by minimizing the gap between the final loss function and the optimal one within limited resources. Finally, extensive numerical results demonstrate the algorithmic advantages in adapting system heterogeneity and admirable performance of the proposed methodologies in speeding up FL 5%–10% and reduce the energy consumption in training about 10% compared with FedProx.
Jing He 0011, Songtao Guo, Dewen Qiao, Lin Yi
IEEE Internet Things J.3
2022 Adaptive Federated Deep Reinforcement Learning for Proactive Content Caching in Edge Computing
abstract
With the aggravation of data explosion and backhaul loads on 5 G edge network, it is difficult for traditional centralized cloud to meet the low latency requirements for content access. The federated learning (FL)-basedproactive contentcaching (FPC) can alleviate the matter by placing content in local cache to achieve fast and repetitive data access while protecting the users’ privacy. However, due to the non-independent and identically distributed (Non-IID) data across the clients and limited edge resources, it is unrealistic for FL to aggregate all participated devices in parallel for model update and adopt the fixed iteration frequency in local training process. To address this issue, we propose a distributed resources-efficient FPC policy to improve the content caching efficiency and reduce the resources consumption. Through theoretical analysis, we first formulate the FPC problem into a stacked autoencoders (SAE) model loss minimization problem while satisfying resources constraint. We then propose an adaptive FPC (AFPC) algorithm combined deep reinforcement learning (DRL) consisting of two mechanisms of client selection and local iterations number decision. Next, we show that when training data are Non-IID, aggregating the model parameters of all participated devices may be not an optimal strategy to improve the FL-based content caching efficiency, and it is more meaningful to adopt adaptive local iteration frequency when resources are limited. Finally, experimental results in three real datasets demonstrate that AFPC can effectively improve cache efficiency up to 38.4$\%$and 6.84$\%$, and save resources up to 47.4$\%$and 35.6$\%$, respectively, compared with traditional multi-armed bandit (MAB)-based and FL-based algorithms.
Dewen Qiao, Songtao Guo, Defang Liu, Saiqin Long, Pengzhan Zhou, Zhetao Li
IEEE Trans. Parallel Distributed Syst.1
2020 Improved evolutionary algorithm and its application in PID controller optimization
Dewen Qiao, Nankun Mu, Xiaofeng Liao 0001, Junqing Le, Fan Yang 0064
Sci. China Inf. Sci.1